SecurityBrief US - Technology news for CISOs & cybersecurity decision-makers
United States
Brackett opens index to test AI agents' effectiveness

Brackett opens index to test AI agents' effectiveness

Mon, 28th Sep 2026 (Today)
Joseph Gabriel Lagonsin
JOSEPH GABRIEL LAGONSIN News Editor

Brackett has open-sourced the Agent Effectiveness Index, a benchmark for assessing how AI agents learn and carry out complex operational tasks. The index publishes initial scores for Brackett, OpenAI's Codex and Anthropic's Claude.

The benchmark is designed to measure agent systems on work that goes beyond factual recall into process execution, exception handling and behavioural retention. Brackett has released the full task set, scoring code and methodology publicly under an MIT licence.

The launch comes as companies test AI agents for operational roles while facing questions over how to judge whether those systems can perform reliably in day-to-day business settings. Many existing evaluations focus on static knowledge or narrow tasks, while businesses often need systems that can apply judgement, stay within approved limits and escalate work to a human when required.

Brackett's framework measures three areas: Business Understanding, Operational Execution and Learning Persistence. In practice, that means testing whether an agent can ground answers in evidence, produce correct outcomes while handling exceptions, and retain newly taught behaviours as processes change.

The first published results compare three agent systems using the same demonstration. Further scoring for execution, transfer and retention will follow as the index develops.

Ehsan Azarnasab, Co-Founder and Chief Scientist at Brackett, described the gap the benchmark is meant to fill.

"Most evaluations for AI agents focus on static knowledge, or what it knows. But this doesn't tell you whether or not that agent can complete the task it was created for, because real world tasks require things like judgement calls, exceptions, or knowing when to bring in a human. These are learned from experiences, not training data. Until now there has been no way to score that capability," said Ehsan Azarnasab, Co-Founder and Chief Scientist at Brackett.

Wider launch

Alongside the benchmark, the San Francisco-based company introduced what it calls its Connected Agentic Workforce platform. The system is intended to turn conversations with workers into software agents that can learn business processes without coding.

Those agents are then linked to one another, to the business systems where they operate and to the people who trained them. Brackett argues that organisations should retain the operational knowledge created through those workflows rather than rely entirely on model providers.

That pitch reflects a broader market debate over whether businesses adopting AI tools are building durable internal know-how or simply renting access to general-purpose models. Start-ups and large technology groups alike have been trying to persuade customers that their tools can capture company-specific practices and decision-making.

Brackett was founded by former technical leaders from Microsoft, Rubrik and Amazon, and is backed by Focal and Heavybit. According to the company, Azarnasab previously worked as Principal Scientist on Microsoft's GenAI Platform team.

The benchmark's initial focus on human demonstration also points to a growing area of interest for AI agent developers. Rather than testing only whether a model can answer a question, firms increasingly want to know whether it can observe a process, reproduce it consistently and adapt when rules or circumstances shift.

That is particularly important in operational settings, where a wrong action can have financial or compliance consequences and where a system that appears plausible may still fail on edge cases. By publishing its task set and methodology, Brackett appears to be inviting outside scrutiny of how those judgements are measured.

Jaideep Sarkar, Co-Founder and Chief Executive Officer of Brackett, said the company viewed the benchmark as part of a larger shift in how businesses work with AI agents.

"The next era of work isn't humans versus agents, it's humans and agents becoming genuinely better together. We built Brackett because we've seen a recurring gap in which companies try to automate tasks, but without a way to build compounding, connected intelligence that they actually own. It's also why we're opening the Agent Effectiveness Index to the world. You can't build a trustworthy agentic workforce without a way to measure it," said Jaideep Sarkar, Co-Founder and Chief Executive Officer of Brackett.