← Research

Evaluation

Your agent improved. Did your evaluation notice?

An October working paper shows how changes in human reviewers can obscure an agent’s progress. Reliable evaluation needs a stable measuring process and evidence that stays current.

This article discusses external research and product documentation. Its illustration is conceptual and it is not an announcement of PersonaSI product functionality.

A flat evaluation chart can look like a verdict: the agent has stopped improving. Before accepting that conclusion, teams should ask whether the people assigning the scores have changed. A new production study makes that question unusually concrete, and offers a useful warning for anyone building a long-lived personal agent.

In a preliminary working paper first posted on October 4, 2026, Liu Zhang and Mark Esposito analyze 2,611 recorded production interviews conducted by a voice-and-video agent. Two reviewers scoring the same batches differed by 0.79 standard deviations on the composite measure. A change in reviewer composition obscured improvement; the reviewer-adjusted, smoothed index rose by 0.53 standard deviations between March and August. Those quantities describe a standardized evaluation index, not percentage-point increases in task success. Reliability of AI Agents, v1 ↗

The study also treats evidence as something that ages. Under its estimated drift, forecast uncertainty reaches the size of the largest observed deployment contrast after roughly five weeks. That estimate is imprecise and specific to one system and period. The authors do not identify a causal benefit from running the evaluation program. This is observational evidence from an interviewing deployment, with assumptions about raters and changing conditions, rather than a universal evaluation schedule. Working paper ↗

The practical implication is to version the measuring process alongside the agent. Preserve the rubric, evaluator identity, task sample, model configuration, tools, and evaluation date. When a reviewer or automated judge changes, score an overlapping set before comparing trends. Otherwise, a release decision may reflect a stricter assessor, easier examples, or a different operating environment rather than a real change in behavior. This is an engineering recommendation drawn from the evidence.

Anthropic’s January 9, 2026 evaluation guide adds another important distinction: finding one successful attempt and succeeding consistently answer different questions. It separates pass@k, which asks whether at least one of several attempts succeeds, from pass^k, which asks whether all succeed. The guide also recommends examining actual outcomes and keeping trials isolated. For an assistant expected to handle the same obligation reliably each week, a best-of-several demonstration is incomplete evidence. Anthropic evaluation guide ↗

OpenAI’s May 29, 2026 guidance on third-party evaluations extends the argument to the surrounding software. Tools, retained state, recovery behavior, and computational budget can change the capability an evaluation reveals. Its recommendations distinguish testing a system’s strongest credible performance from comparing systems under equivalent conditions. Reports should state which claim they support and disclose the relevant setup. A higher score earned with a different toolset or much larger budget can still be useful, but readers need those conditions to interpret it. OpenAI evaluation guidance ↗

Human interaction deserves its own measurements. HAS-Bench, first posted on July 5, 2026, explicitly evaluates clarification, use of feedback, control, and interaction cost alongside task outcomes. It provides a research framework for varying human participation rather than treating every intervention as equivalent. Its experiments use simulated users, so their findings should not be read as a forecast of behavior across all real people. HAS-Bench, v1 ↗

For PersonaSI research, that suggests separating several questions: Did the task finish? Did the agent honor a correction? Did it ask a useful question? Did it require avoidable supervision? An aggregate success score can conceal a system that completes more work by consuming more human attention.

A useful release review would combine repeated task trials, unchanged reference cases, newly observed failures, and a short account of what changed in the evaluator. It would also identify who can fix a failure once the evaluation exposes it. Confidence should be attached to a defined claim, a tested configuration, and a date. These are proposed research and evaluation practices, not claims about capabilities currently implemented in PersonaSI.

Sources