@openbenchmarksi
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- Web
Account-level information from X, not a live location or the device used for a specific post.
Independent benchmarks to help you discover, compare, and choose APIs and agents for your use cases.
San Francisco, CA
Joined July 2026
- Tweets63
- Following42
- Followers58
- Likes35
Introducing LegalParseQA Benchmark:
An independent test of how well document parsers handle heavily redlined contracts, and whether an AI agent can still tell what was deleted.
@AnthropicAI @OpenAI @llama_index @reductoai @datalabto @ExtendHQ @Pulse__AI @MistralAI
There's no single best document parser. There's a best one for your file type, latency budget, cost per correct ans.
Benchmark runner, scorer, a public sample dataset are open, so results can be checked locally. Contracts are built from RedlineBench by Crosby Legal (CC-BY-4.0).
Eval harness is Open Source
Eval harness + code: github.com/openbenchmarks-la…
Public dataset: huggingface.co/openbenchmark…
Full benchmark: openbenchmarks.com/document-…
Benchmark Insights: Company Enrichment
Metric: Field accuracy
Every time you enrich a company record, you're writing new values into your CRM. And those values end up deciding who gets the lead, which territory it lands in, how it's scored and which segment it belongs to.
Across 11 endpoints, field accuracy ranges from 70.4% to 92.6%. And 8 sit between 85% and 93%, so at least 4 in 5 returned values are right. @useapolloio (92.6%), @PeopleDataLabs (91.2%), @p0 (89.4%), @nimble_search (lite, 88.1%) and @ExploriumAI (87.6%) lead the pack.
Eval harness is Open Source
Eval harness + code: github.com/openbenchmarks-la…
Public dataset: huggingface.co/datasets/open…
Full benchmark: openbenchmarks.com/company-e…
The agent reads, decides what to search or fetch next, and the context grows. Median task cost captures both sides: the search and fetch fees, plus the LLM tokens spent turning results into an answer.
We looked at this across 2 benchmarks:
Low cost doesn't come at the expense of results. On Hard Retrieval, Search & Fetch, @Tiny_Fish completes 79.0% of tasks, uses fewest tokens of any endpoint (a median of 12,844 per task) and averages 1.32s per search.
Eval harness is Open Source
Eval harness + code: github.com/openbenchmarks-la…
Public dataset: huggingface.co/datasets/open…
Full benchmark: openbenchmarks.com/web-searc…
Benchmark Spotlight: @Tiny_Fish
Benchmarks: Websearch
- Hard Retrieval
- Multi-hop Search
Metric: Median Task Cost
When an agent uses a search API, every result it returns lands in the model's context, and the model pays to read it.