We Put Echosaw on the Bench. Here's What the Numbers Said.
We built five test benches for AI media tools, entered Echosaw in two, and published everything - harness, inputs, scoring code, and results. The bench put us first in both - and the tables still gave us a to-do list.
Last week we launched Curated Lists: five measured comparisons of tools in the categories we live in. Speech-to-text APIs. Multimodal embedding models. S3-compatible object storage. AI video transcript and summary services. AI video search tools that answer questions about your media.
Three of those categories are things we buy. Two are things we sell. We entered Echosaw in the two we sell. Five lists is the whole set: we picked the categories before a single number existed, ran five, and published five. There is no pile of comparisons we quietly lost and left in a drawer. The order on each page is set by a fixed rubric applied to measured results, the same rubric for every entry, and Echosaw is tagged as ours wherever it appears. We did not pick our position; the bench did.
The rules everyone followed
Every entrant in a list got the same treatment: the same input file, the same scoring script, the same reference transcript or question set, the same rubric weights, and one run on the test date. Echosaw got no other treatment. The order on the page is what falls out of those numbers.
A vendor-authored comparison is worth nothing unless you can check it, so all of that is public. Each measured list links to a harness page on echosaw.com with the input media, the scoring code, the reference data, and the raw JSON every service returned. If we say a competitor got 10 of 14 answers right, you can open the file and count.
Where we could not run a service - no free tier and no trial credits we could test with - we said so and placed it on documentation only, with no numbers in the results table. Google Video Intelligence, Azure Video Indexer, NotebookLM, Mixpeek, and Fireflies are all in that bucket. We would rather show an empty cell than a guessed one.
Everything was run on the same 10-minute clip of a four-person meeting from the AMI corpus, recorded on a room camera. It is hard audio: far-field, overlapping talkers, a whiteboard that nobody is facing. One clip is not a benchmark. It is enough to show real differences and not enough to generalise from, and the pages say that too.
The feel-good part
The bench put Echosaw first in both lists it entered.
On transcript and summary services, our summary got all three financial figures in the meeting right - the 25 euro price, the 50 million euro target, the 12.50 euro cost cap - and was the only service that described what was on screen, because we look at the video track and not just the audio. Otter got the revenue target wrong by a factor of three. Twelve Labs skipped the figures entirely.
On AI video search tools, we asked each tool 14 questions about the meeting. Echosaw returned the right file first for 14 of 14 - out of a library of 17 unrelated items - and landed inside the gold time span for 13 of 14. Asked the question in words with no hint about which file to look in, POST /v1/media/ask answered 12 of 14 correctly and deep-linked each answer to the moment it was spoken. Twelve Labs, searching an index that contained only the one video, put the right clip first for 4 of 14.
We enjoyed that. But the work continues.
Back to work
Here is what the lists put on our board, in order:
- Transcript accuracy. A two-point WER gain on meeting audio is the largest single improvement available to us and we feel worth it.
- The "price" collision. A question about a price in a recording should be treated as a question about the recording. We are looking at how intent is routed.
- Turnaround. More compute, more parallelism, fewer passes waiting on each other. 178.8 seconds has to come down without loss of quality.
- Search latency. 2.55 seconds is fine for a person and slow for a program. We would like it under a second.
Why bother
We could have written five "top tools" articles with a paragraph per vendor and put Echosaw quietly at the top without full transparency. We chose transparency.
We would rather publish a number we do not love next to a competitor's number that beats it, with the file that produced both. It is the same bargain we ask our users to accept every day: don't take our word for it, here's the evidence. It would be strange to sell that and not practise it.
Read the lists. Open a harness. Run the harness. Count for yourself.
Every curated list publishes everything behind it: the harness, the input media, the scoring code, the reference data, and the full results from every service tested. Echosaw is tagged as ours wherever it appears.