Openbenchmarks
Independent benchmarks for build-vs-buy decisions. Verifiable, reproducible and open source.
Last run on 3 Oct 2026Benchmarks
Web Search for AI
Web Search APIs benchmarked across three tasks: factual lookup, hard retrieval, and multi-hop search.
- Most accurate
Exa fast99.3%
- Fastest
Parallel turbo348 ms- Lowest cost
TinyFishFree
- Most tasks
Perplexity77.3%- Fastest
Perplexity18.3 s- Lowest cost
TinyFish$0.035
- Most tasks
Exa deep83.0%
- Fastest
Perplexity21.5 s- Lowest cost
TinyFish$0.050
- Leads F1
Parallel basic46.5- Fastest
Brave Search43.5 s- Lowest cost
TinyFish$0.212
- Leads F1
Exa deep48.2
- Fastest
Brave Search45.1 s- Lowest cost
TinyFish$0.219
Multi Turn Company Search
Search APIs tested as tools for a research agent finding complete company sets across three- and four-part constraints. Ranked on precision, recall, F1, exact-set accuracy, latency, and cost.
- Leads F1
Parallel basic46.5- Fastest
Brave Search43.5 s- Lowest cost
TinyFish$0.212
- Leads F1
Exa deep48.2
- Fastest
Brave Search45.1 s- Lowest cost
TinyFish$0.219
Company News
Which API answers questions about recent company events. Web search APIs against dedicated news indexes, same questions.
- Most accurate
Exa (type=fast)99.3%
- Fastest
Parallel (mode=turbo)348 ms- Lowest cost
TinyFishFree
Web Search for Coding Agents
Which search API helps a coding agent find an opaque detail in official Enterprise SaaS docs and ground the patch in a URL it actually retrieved.
- Most tasks
Perplexity77.3%- Fastest
Perplexity18.3 s- Lowest cost
TinyFish$0.035
- Most tasks
Exa deep83.0%
- Fastest
Perplexity21.5 s- Lowest cost
TinyFish$0.050
Document Processing Benchmark
Specialised document parsers benchmarked on heavily redlined contracts, delivered as text-layer PDFs and as scanned images.
- Most accurate
LlamaParse agentic80.0%- Fastest
Reducto7s- Parser $ / 1k correct
Mistral OCR$14.04
- Most accurate
Datalab track changes79.7%- Fastest
Reducto8s- Parser $ / 1k correct
Datalab track changes$19.13
Company Lookalikes
Tests how well company-search and lookalike APIs turn a seed company domain or description into a ranked list of genuinely similar businesses. Use it to compare tools for account discovery, prospecting, market mapping, and TAM expansion.
- Most precise
Seltz63.4%
- Fastest
Discolike1.2 s- Relevant / $1
Seltz12,683
Company Firmographic Enrichment
282 company domains × 11 APIs, normalized into the same seven-field active scoring contract. Compare enrichment success rate, field accuracy, coverage, company match rate, latency, and cost by workflow.
- Correct field yield
People Data Labs89.0%- Fastest
People Data Labs274 ms- Lowest cost
Nimble lite$0.31
Company Funding Data Enrichment
17 funding vendors, judged on latest funding stage against human labelled ground truth.
- Correct stage yield
Firecrawl92.3%- Fastest
People Data Labs276 ms- Lowest cost
Ocean.io$1.35
- Correct stage yield
Firecrawl100.0%- Fastest
Apollo320 ms- Lowest cost
Nimble lite$0.23
Inference Benchmark
Serverless inference providers serve the same GLM 5.3 Flash model on 600 latency-sensitive lookup, classification, and extraction requests each. Compare response latency, streaming speed, task success, and failures without collapsing them into one overall rank.
- Fastest P99
Baseten3.2 s
- Most reliable
Fireworks AI99.5%- Fastest streaming
Baseten240 tok/s
Voice Agents Latency Benchmark
TTFAB — how long a caller waits before a voice AI agent starts speaking — measured over real phone calls from the call's own audio, never platform-reported timestamps.
- Fastest median
Telnyx1,296 ms- Fastest P95
ElevenLabs1,768 ms- Lowest cost
Telnyx$0.0500/min
Text-to-Speech
32 text-to-speech models on Time to First Audio (TTFA) + Word Error Rate, measured by Coval under production-realistic conditions. Mirrored with attribution.
- Lowest WER
Cartesia sonic-3.61.6%- Fastest
Fluxions vui45 ms
Live Speech-to-Text
30 live speech-to-text models on Word Error Rate and Time to Final Segment (TTFS), measured by Coval under production-realistic conditions. Rolling 7-day window, mirrored with attribution. Ranked here by WER, not median TTFS.
- Lowest WER
AssemblyAI universal-3.6-pro2.1%- Fastest
Baseten qwen3-asr-1.7b31 ms
Structured Speech-to-Text
300 human-recorded workplace utterances × 17 ASR systems, checked for exact recovery of 1,482 structured values — emails, phone numbers, CLI flags, file paths, IDs. Cell value is Task Success Rate — recordings with every value correct.
- Task success
Deepgram Nova-371.0%- Value accuracy
ElevenLabs Scribe v291.8%- Lowest WER
AssemblyAI Universal 3.5 Pro3.9%
- Task success
Deepgram Nova-361.0%- Value accuracy
OpenAI GPT Realtime (Whisper)88.7%
- Lowest WER
ElevenLabs Scribe v2 Realtime5.1%