@openbenchmarks

Independent benchmarks to help you discover, compare, and choose APIs and agents for your use cases.

San Francisco, CA
Joined July 2026
Introducing LegalParseQA Benchmark: An independent test of how well document parsers handle heavily redlined contracts, and whether an AI agent can still tell what was deleted. @AnthropicAI @OpenAI @llama_index @reductoai @datalabto @ExtendHQ @Pulse__AI @MistralAI
2
1
3
8
270
There's no single best document parser. There's a best one for your file type, latency budget, cost per correct ans. Benchmark runner, scorer, a public sample dataset are open, so results can be checked locally. Contracts are built from RedlineBench by Crosby Legal (CC-BY-4.0).
1
31
Benchmark Insights: Company Enrichment Metric: Field accuracy Every time you enrich a company record, you're writing new values into your CRM. And those values end up deciding who gets the lead, which territory it lands in, how it's scored and which segment it belongs to.
2
3
5
117
Across 11 endpoints, field accuracy ranges from 70.4% to 92.6%. And 8 sit between 85% and 93%, so at least 4 in 5 returned values are right. @useapolloio (92.6%), @PeopleDataLabs (91.2%), @p0 (89.4%), @nimble_search (lite, 88.1%) and @ExploriumAI (87.6%) lead the pack.
1
1
24
The agent reads, decides what to search or fetch next, and the context grows. Median task cost captures both sides: the search and fetch fees, plus the LLM tokens spent turning results into an answer. We looked at this across 2 benchmarks:
1
2
55
Low cost doesn't come at the expense of results. On Hard Retrieval, Search & Fetch, @Tiny_Fish completes 79.0% of tasks, uses fewest tokens of any endpoint (a median of 12,844 per task) and averages 1.32s per search.
1
1
20
Benchmark Spotlight: @Tiny_Fish Benchmarks: Websearch - Hard Retrieval - Multi-hop Search Metric: Median Task Cost When an agent uses a search API, every result it returns lands in the model's context, and the model pays to read it.
3
2
11
285