CORTEXA
← Browse
arxivcs.AIcs.HC2026-07-16

Project Kaleidoscope: Contextual, Human-Aligned Evaluation for Real-World AI Applications

Leanne Tan, Rohan Jaggi, Shaun Khoo, Roy Ka-Wei Lee

Evaluations (Evals) are a deployment bottleneck for real-world AI applications: public benchmarks rarely match a team's users, context, or policies, and human review is often tedious to scale. Motivated by our work with AI applications in the public sector, this project addresses recurring evaluation challenges encountered when applications must satisfy local policy and governance requirements. We present Kaleidoscope, an integrated workflow for contextual functional evaluation that links persona-based test generation, contextualized rubrics, and human review for reliability-gated automated scoring. Generated test cases are scored against application-specific rubrics; human annotations provide reviewable labels; and LLM judges automate scoring only when their agreement with those labels meets a configured threshold. Kaleidoscope is therefore a practical, inspectable, iterative workflow for product teams. We report early evidence from a three-week pilot across four organizational use cases and custom-rubric judge experiments on 108 annotated Q\&A pairs spanning four domains and 14 evaluation dimensions. The results highlight useful features for end-to-end reliable, automated scoring.

View free PDFSource page

Related papers

arxivcs.CLcs.AIcs.DLcs.HC2026-07-18

Real-World Evaluation of an AI Agent Drafting Translational Impact Summaries

Mohammad Arvan, Amber E. Osterholt, Bailee Rue, Yuvaneswaren R. Sureshbabu, Krishna R. Patel, Rebecca T. Feinstein, et al.

Introduction. Clinical and Translational Science Award (CTSA) programs must document their scholars' research impact, but assembling each scholar's record by hand takes staff an estimated 15 hours and does not scale to a full cohort. An artificial intelligence (AI) agent could se…

View free PDFSource page
arxivcs.CVcs.AIcs.HC2026-07-09

VEGAS: Human-Aligned Video Caption Evaluation via Gaze

Shenghui Chen, Po-han Li, Ximeng Sun, Shijia Yang, Emad Barsoum, Zicheng Liu, et al.

Vision-language models excel at video captioning, yet typically generate descriptions that fail to capture individual viewers' attention. We propose VEGAS (Video caption Evaluation via GAze Score), a training-free metric that leverages test-time gaze to sample personalized, atten…

View free PDFSource page
arxivcs.HCcs.AIcs.CY2026-06-28

LLMography: Transforming Human-AI Conversations into Traceability, Oversight, and Auditability Indicators

Mohammed Bousmah

The growing use of Large Language Models (LLMs) in education, software engineering, academic writing, and technical documentation raises a key question: how can we evaluate not only AI-assisted outputs, but also the interaction process that produced them? Current debates often fo…

View free PDFSource page
arxivcs.AIcs.CEcs.HC2026-07-14

Networked Intelligence: Active Shared Context Graphs for Human-AI Team Science

Sutanay Choudhury, Jeffrey J. Czajka, Lummy M. O. Monteiro, Erin Bredeweg, Jason McDermott, Katherine Wolf, et al.

Most AI-for-science systems focus on scaling a single reasoning process by using better models, larger context windows, long-horizon agentic execution, or digital co-scientists working with one principal user. However, challenging scientific problems are rarely solved by one reas…

View free PDFSource page
arxivcs.HCcs.AI2026-07-02

What Types of Human-AI Teams Exist?

Nathan Hughes, Ibrahim Habli

Human-AI teaming has received increasing attention in the literature. However, the range of studies conducted in multiple domains make it difficult to understand what types of teams are being studied, and in what ways are they similar/different from one another. In this study, we…

View free PDFSource page
arxivcs.AIcs.HC2026-07-08

Learning social norms enhances compatibility in dynamic human-AI coordination

Yi Yang, Siyuan Liu, Xin Gao, Huamu Sun, Chao Liu, Qing Zhou, et al.

Humans continuously coordinate with others in dynamic interactions, often through implicit, hard-to-quantify social norms that act as shared tacit expectations among interacting agents. As AI agents, including large language models (LLMs), become embedded in daily life, they incr…

View free PDFSource page