Media Intelligence
Best AI Video Transcript & Summary Services in 2026: Measured on One Meeting Video
Every service here takes a video file and gives back a transcript and a summary. The question is how accurate the transcript is, whether the summary gets the facts right, how long you wait, and what it costs. We uploaded the same 10-minute meeting recording to each service we could reach, scored the transcript against a human reference, and checked the summary against a fixed list of facts stated in the meeting.
Echosaw is an entry in this list. We built the harness, ran it, and publish every raw output below. The ranking follows the rubric, not the byline.
The test video is one 10-minute sample of a four-person meeting recorded on a room camera. It is a hard case for transcription (overlapping talkers, far-field audio) and an easy case for summarisation (one topic, clear structure). It is not a universal measure: a lecture, a product demo, or a screen recording would rank these services differently.
- Tested
- September 10, 2026
- Last updated
- September 10, 2026
- Tested by
- Devin (AI engineer, Echosaw)
- Reviewed by
- Matthew Simpson
Disclosure: Echosaw is our product and appears in this list. It is scored on the same rubric and dataset as every other entry.
Quick answer
- 1
Best for teams that want transcript, summary, timeline, and visual analysis from one API call, then search across the results.
- 2ScreenAppMeasured
Best for individuals who want a chaptered summary of a recording fast.
- 3OtterMeasured
Best for meeting notes from Zoom, Teams, and Meet where a bot joins the call; file import is secondary.
- 4Twelve LabsMeasured
Best for developers building video search or open-ended video Q&A who want a prompt-driven summary rather than a fixed report.
- 5Azure AI Video IndexerDoc-verified
Best for azure-committed organisations that need transcription, translation, and visual insights under one Azure resource, including edge deployment via Arc.
- 6Google Cloud Video IntelligenceDoc-verified
Best for GCP pipelines that need labels, shots, and text detection alongside a transcript and will build their own summary step.
Measured results
4 of 6 entries were run on the same input. Entries marked doc-verified were not run and are excluded from this table.
| Tool | Version tested | Transcript WER | Summary facts | Turnaround | Summary length |
|---|---|---|---|---|---|
| EchosawOur product | Echosaw gamma pipeline, POST /v1/analyze/url | 19.3%34 sub / 146 del / 52 ins | 9 / 12all three financial figures correct | 178.8 ssingle analysis pass, transcript + summary + visual | 122 words |
| ScreenApp | ScreenApp web app, free plan; transcript provider "xai", summary provider "xai" as reported by the app | 20.2%48 sub / 134 del / 61 ins | 12 / 12six chapters + overview | 60.5 supload start to summary generatedAt | 507 words |
| Otter | Otter web app, Basic plan, English (US) file import | 18.2%43 sub / 129 del / 46 ins | 6 / 12revenue target wrong (15M vs 50M) | 74 screated to modified timestamps | 207 words |
| Twelve Labs | marengo3.0 index (visual+audio) transcription + pegasus1.5 analyze, twelvelabs SDK 1.3.4 | 30.7%63 sub / 165 del / 140 ins (duplicated words) | 8 / 12no financial figures | 125.8 s3.5 upload + 100.8 index + 21.4 analyze | 99 words |
How we evaluated
The video is the first 600 seconds of AMI meeting ES2002a (corner-camera view with the meeting-room audio track, H.264/AAC). The original 352x288 file was rejected by Twelve Labs as too low a resolution, so every service received the same 640x480 upscale (ffmpeg, libx264 CRF 20, audio stream copied). Each service was run once, on the test date, from a free or free-tier account created with an echosaw.com address; Echosaw ran on our own production API.
Transcript WER is computed with jiwer against the 1,200-word AMI reference transcript after both texts pass through the Whisper English normaliser (lowercasing, punctuation, number and contraction normalisation, "um"/"uh" dropped), identical to the speech-to-text list. Summary facts is the number of twelve facts stated in the meeting (purpose, brief, the four roles, the three-stage process, the whiteboard icebreaker, the 25 euro price, the 50 million euro target, the 12.50 euro cost cap, international considerations) that appear in the service's own summary, matched by pattern in score.py. It measures coverage, not writing quality.
Turnaround is wall-clock seconds from the start of upload to summary available. For Twelve Labs and Echosaw it is measured by the runner; for ScreenApp it is the noted upload start to the summary's generatedAt timestamp; for Otter it is the conversation's created to last-modified timestamps as reported by its API. ScreenApp and Otter have no public API on their free plans, so their transcript and summary were captured from the JSON their web apps load and fed to the runner with --import; the capture steps are documented in run.py.
Pricing is the list price on each vendor's public pricing page on the test date. Google Video Intelligence and Azure Video Indexer were not run (they require a billed GCP project or Azure subscription that we did not set up for this round); they are ranked on published capability and price only and do not appear in the results table.
Dataset
AMI ES2002a, first 600 s, corner camera (640x480 upscale) — Kick-off meeting of a four-person design team: introductions, a whiteboard icebreaker, project finance, and international considerations. Reference transcript (1,200 words) built from the AMI word-level annotations by the speech-to-text harness's build_dataset.py. License: AMI Meeting Corpus CC BY 4.0.
Harness and raw output: /curated-lists/harness/video/
Ranking rubric
- Transcript accuracy35%
- Word error rate against the human reference on the meeting video.
- Summary quality30%
- Facts from the meeting correctly present in the summary, and absence of wrong facts.
- Turnaround and workflow15%
- Time from upload to summary, and whether the result is reachable by API.
- Price20%
- List price for one 10-minute video on the cheapest plan that supports the workflow.
The rankings
Orange Sky Software · Best for teams that want transcript, summary, timeline, and visual analysis from one API call, then search across the results
Got all three financial figures right and was the only service to describe what was on screen (the room, the whiteboard), at the cost of the slowest turnaround.
Strengths
- 19.3% transcript WER, within about a point of the best result on this audio.
- Summary covered 9 of 12 facts and stated all three financial figures correctly; it also identified the room and the whiteboard activity from the video track, which no audio-only service can do.
- Single POST /v1/analyze/url returns transcript, summary, timeline, key phrases and billing in one report; no separate index step.
Limitations
- 178.8 s turnaround for a 10-minute video, the slowest measured: the pipeline runs visual analysis and moderation whether or not you need them.
- Summary omitted the "original, trendy, user-friendly" brief and the three-stage process, and did not give the industrial designer's role.
- $0.43/min audio+video on the Agency plan ($4.30 for this video) is the highest per-minute rate in the test; the Starter plan is $0.58/min. Audio-only analysis is $0.22-0.37/min.
Pricing: Subscription $9-49/month plus usage: audio+video $0.58/min (Starter) to $0.43/min (Agency); audio-only $0.37 to $0.22/min
ScreenApp Pty Ltd · Best for individuals who want a chaptered summary of a recording fast
Fastest turnaround and the most complete summary in the test (12 of 12 facts, six chapters), but the free plan caps you at three recordings and there is no API below the Business plan.
Strengths
- 60.5 s from upload to summary, the fastest measured.
- Summary hit all 12 facts, including the correct 50 million euro target, with chapters and speaker attribution.
- Free plan needs no card and includes the full transcript; a 24.9 MB upload processed on the free tier.
Limitations
- 20.2% transcript WER, mid-pack; the transcript rendered one participant's name two different ways.
- Free plan allows 3 recordings total and 1 transcription a month; transcript download and export are paid features (we captured ours from the app's own JSON).
- API access requires the $34/month Business plan.
Pricing: Free (3 recordings, no card); Growth $19/month billed annually ($228/year); Business $34/month billed annually with API access
Otter.ai · Best for meeting notes from Zoom, Teams, and Meet where a bot joins the call; file import is secondary
Best transcript accuracy in the test and a fast turnaround, but the summary misstated the revenue target and the free plan allows only three file imports for the life of the account.
Strengths
- 18.2% transcript WER, the lowest measured on this meeting audio.
- 74 s from upload complete to processing done; summary, outline, and keyword list generated automatically.
- Basic plan is free with 300 transcription minutes a month and no card.
Limitations
- Summary stated the revenue goal as 15 million euro; the recording says fifty million. It also named participants without their roles, so it hit 6 of 12 facts.
- Three lifetime file imports on Basic; Pro ($16.99/user/month) allows 10 imports a month.
- No self-serve API; results were captured from the web app.
Pricing: Basic free (300 min/month, 3 lifetime imports); Pro $16.99/user/month or $8.49 billed annually; Business $30/user/month or $24 annually
Twelve Labs · Best for developers building video search or open-ended video Q&A who want a prompt-driven summary rather than a fixed report
A developer API with a generous free tier and a prompt you control, but the transcript returned by the index carried duplicated words that doubled its error rate.
Strengths
- Pegasus summary was accurate on everything it stated (roles, brief, icebreaker, international sales) and correctly noted no decisions were taken.
- Free plan: 600 minutes of indexing, no card; indexes persist 90 days.
- Summary prompt is yours; the same indexed video can then be searched with Marengo.
Limitations
- 30.7% transcript WER: the transcription attached to the index repeated many words across adjacent segments ("this this is is the the kickoff kickoff"), producing 140 insertions. Speech transcription is a by-product of indexing, not a first-class output.
- Summary left out all three financial figures and the three-stage process (8 of 12 facts).
- Two-step workflow (index, then analyze); 125.8 s total of which 100.8 s was indexing.
Pricing: Free 600 minutes; Developer pay-as-you-go: Marengo indexing $0.042/min, Pegasus analyze $0.0292/min input + $0.0075 per 1k output tokens (about $0.72 for this video)
Microsoft Azure · Best for azure-committed organisations that need transcription, translation, and visual insights under one Azure resource, including edge deployment via Arc
The broadest insight catalogue of any service here (speaker indexing, OCR, faces, topics, generative summaries) and the only one deployable on-premises; not measured in this round.
Strengths
- Transcription with speaker indexing, translation, topics, named entities, and generative textual summaries per the product documentation.
- Runs in the cloud or on Arc-enabled Kubernetes for data-residency requirements.
- Trial account: up to 10 hours of free indexing for website users and 40 hours for API users, per the pricing page.
Limitations
- Not run in this test; WER, summary coverage, and turnaround figures are unavailable.
- Pricing is per input minute by preset (Basic, Standard, Advanced audio and video); the rates are shown only after selecting a region on the pricing page, so no single list price is quoted here.
- Requires an Azure subscription and a Video Indexer account bound to a storage account before the first upload.
Pricing: Per input minute by preset (Basic / Standard / Advanced audio and video); rates vary by region. Trial: 10 h free (web), 40 h free (API)
Google Cloud · Best for GCP pipelines that need labels, shots, and text detection alongside a transcript and will build their own summary step
A feature-per-minute annotation API with 1,000 free minutes a month of each feature; it returns a transcript but no summary, so it is a component rather than a complete service.
Strengths
- Speech transcription $0.048/min after 1,000 free minutes per month per feature, per the pricing page.
- Label, shot, explicit-content, object, text, logo, face, and person detection in the same request.
Limitations
- Not run in this test; WER and turnaround figures are unavailable.
- No summarisation feature: producing a summary requires a second call to a language model (e.g. Gemini on Vertex AI), which is priced separately.
- Requires a GCP project with billing enabled even to use the free tier.
Pricing: Per feature per minute after 1,000 free minutes/month: speech transcription $0.048, label detection $0.10, shot detection $0.05, object tracking $0.15
Frequently asked questions
- Which service produced the most accurate transcript?
- On our 10-minute AMI meeting video, Otter had the lowest word error rate (18.2%), followed by Echosaw (19.3%) and ScreenApp (20.2%). Twelve Labs was 30.7% because the transcription attached to its index repeated words across segment boundaries. All four are within the range we measured for dedicated speech-to-text APIs on the same audio (16-33%).
- Which service produced the best summary?
- ScreenApp covered all 12 facts we checked, with chapters and speaker names. Echosaw covered 9 and was the only service to describe the room and whiteboard from the video track. Twelve Labs covered 8 with no errors but omitted every number. Otter covered 6 and misstated the revenue target as 15 million euro instead of 50 million, the only factual error in any summary.
- Why is Echosaw ranked first when it was not the most accurate transcript?
- The rubric weights transcript accuracy 35%, summary quality 30%, turnaround and workflow 15%, and price 20%. Echosaw was within 1.1 points of the best WER, got every number in the summary right, and is the only measured service with a self-serve API. It lost on turnaround (slowest) and price (highest per minute). ScreenApp had the better summary and speed but no API below its Business plan. Echosaw is our product; the raw outputs are published so you can weigh the rubric differently.
- Otter alternative: which services did better on our meeting video?
- Otter had the best transcript (18.2% WER) but the weakest summary (6 of 12 facts, with the revenue target misstated as 15 million euro instead of 50 million), and its free plan allows three file imports for life with no self-serve API. ScreenApp covered all 12 facts in 60.5 s at 20.2% WER, and Echosaw covered 9 of 12 with every financial figure correct at 19.3% WER and is the only measured service with a self-serve API. Twelve Labs had a worse transcript (30.7% WER) but a better summary (8 of 12 facts, no errors) and a free 600-minute API tier.
- Why were Google and Azure not tested?
- Google Video Intelligence needs a GCP project with billing enabled and Azure Video Indexer needs an Azure subscription; neither offers a card-free path. We did not set those up for this round, so both are ranked on documented capability and price and marked documentation-verified.
- How can I reproduce these results?
- The runner (run.py), scorer (score.py), and the raw transcript, summary, and timing JSON for each service are published at echosaw.com/curated-lists/harness/video/. Twelve Labs and Echosaw run from API keys you supply as environment variables; ScreenApp and Otter have no free-plan API, so run.py accepts the JSON their web apps load via --import. The video is the first 600 seconds of AMI ES2002a, upscaled to 640x480 with the ffmpeg command in the methodology.