Chrissa Automates logoChrissa Automates

Original research

Tests and experiments built around real observations

This section is reserved for first-hand experiments, benchmarks, and structured tests. Research is only published here when the methodology, scope, findings, and limitations can be documented without inventing data.

CAutomates Agent Utility Index

What the benchmark series is for

The Agent Utility Index tests tools that AI agents depend on and measures the cost of getting usable work done, not only the advertised cost of an API request. It is intended for AI engineers, agent builders, technical founders, and automation teams comparing retrieval, search, browser, enrichment, extraction, and research infrastructure.

Task-level economics

Measure cost per correct result, retries, token burden, evidence quality, and downstream usability instead of request price alone.

Provider selection

See where providers overlap, where difficult pages separate them, and how conclusions change on a fresh sample.

Routing research

Test if a multi-provider agent can beat a strong single provider before adding production routing complexity.