Research preprint; authors report PALM Workshop acceptance2026-10-06
When a plan becomes a false memory
arXiv
Source: AgentMemGate: Addressing Speculation Contamination in Conversational Assistant Memory
Authors: Chirag Sharma; Benjamin Fowlersmith; Karime Maamari
AgentMemGate addresses a specific failure in conversational memory: a tentative plan can be stored as a fact about the user. The proposed approach keeps prospective changes separate until later evidence confirms or abandons them. Controlled experiments suggest that testing intermediate memory states can reveal errors that a final-memory check misses.
Why it matters: A personal agent should distinguish what a user is considering from what has actually happened. This is relevant to editable memory, source provenance and explicit state transitions.
Boundary: The evaluation uses constructed, LLM-generated conversations, small datasets and one fixed profile field per conversation. It does not establish general deployment safety, and effects on queries about unresolved plans remain untested. We have not independently reproduced the results.
Read source ↗Accessible full text ↗Official AI lab blog2026-09-10
Introducing the Agents API
OpenAI
OpenAI presents a managed harness for long-running agents, tool discovery, sandboxes and delegated subagents. The useful PersonaSI signal is architectural: agent quality depends on the surrounding execution, context and recovery system—not only the model.
Boundary: Product announcement; customer performance figures are attributed examples, not independent comparative evidence.
Read source ↗arXiv; forthcoming Mensch und Computer 2026 proceedings2026-08-24
A Framework of User Experience Principles for Human-AI Agent Interaction in the Workplace
SAP / University of Missouri / Hochschule Fresenius
The paper derives eight UX principles for workplace agents through workshops, expert review, meta-analysis and interviews. It strengthens the case for treating controllability, transparency and collaboration as measurable interaction requirements rather than decorative trust language.
Boundary: Exploratory framework and foundation for future empirical validation; not proof that one interface pattern works across every workplace.
Read source ↗Official engineering blog2026-01-09
Demystifying evals for AI agents
Anthropic
Anthropic argues that evaluation suites turn ambiguous product expectations into explicit behavior tests and make model upgrades easier to assess. For specialist agents, a stable bank of representative tasks can connect corrections to measurable regressions and improvements.
Boundary: Engineering guidance and company examples; teams still need domain-specific human validation and realistic task design.
Read source ↗Independent research organization2026-09-08
Task-Completion Time Horizons of Frontier AI Models
METR
METR estimates the difficulty of tasks agents can complete by comparing success probability with the time a human expert needs. Its most important editorial warning is that a time horizon is not the duration an agent can operate autonomously and is heavily shaped by the task suite and agent setup.
Boundary: The suite concentrates on software engineering, machine learning and cybersecurity; results do not generalize to all valuable knowledge work.
Read source ↗Work Trend Index; includes LinkedIn leadership research2026-05-05
Agents, human agency, and the opportunity for every organization
Microsoft WorkLab
Microsoft’s report combines survey data, productivity signals and expert interviews to argue that agent adoption depends on how organizations redesign work, management and learning—not merely on individual tool use. This aligns with PersonaSI’s interest in preserving human judgment while delegating execution.
Boundary: Vendor research with broad organizational claims; reported associations should not be read as proof that agents cause productivity gains.
Read source ↗Independent academic index2026-04-13
The 2026 AI Index Report
Stanford Institute for Human-Centered AI
Stanford’s 2026 index describes a widening gap between technical capability and the institutions needed to govern, evaluate and understand AI. For PersonaSI, the relevant signal is the growing value of documented, independent measurement as deployment accelerates and transparency declines.
Boundary: A wide field-level report rather than an evaluation of PersonaSI or a single agent architecture.
Read source ↗