← All curated lists

Media Intelligence

Best AI Video Transcript & Summary Services in 2026: Measured on One Meeting Video

Every service here takes a video file and gives back a transcript and a summary. The question is how accurate the transcript is, whether the summary gets the facts right, how long you wait, and what it costs. We uploaded the same 10-minute meeting recording to each service we could reach, scored the transcript against a human reference, and checked the summary against a fixed list of facts stated in the meeting.

Echosaw is an entry in this list. We built the harness, ran it, and publish every raw output below. The ranking follows the rubric, not the byline.

The test video is one 10-minute sample of a four-person meeting recorded on a room camera. It is a hard case for transcription (overlapping talkers, far-field audio) and an easy case for summarisation (one topic, clear structure). It is not a universal measure: a lecture, a product demo, or a screen recording would rank these services differently.

Tested
September 10, 2026
Last updated
September 10, 2026
Tested by
Devin (AI engineer, Echosaw)
Reviewed by
Matthew Simpson

Disclosure: Echosaw is our product and appears in this list. It is scored on the same rubric and dataset as every other entry.

Quick answer

  1. 1
    EchosawOur productMeasured

    Best for teams that want transcript, summary, timeline, and visual analysis from one API call, then search across the results.

  2. 2
    ScreenAppMeasured

    Best for individuals who want a chaptered summary of a recording fast.

  3. 3
    OtterMeasured

    Best for meeting notes from Zoom, Teams, and Meet where a bot joins the call; file import is secondary.

  4. 4
    Twelve LabsMeasured

    Best for developers building video search or open-ended video Q&A who want a prompt-driven summary rather than a fixed report.

  5. 5

    Best for azure-committed organisations that need transcription, translation, and visual insights under one Azure resource, including edge deployment via Arc.

  6. 6

    Best for GCP pipelines that need labels, shots, and text detection alongside a transcript and will build their own summary step.

Measured results

4 of 6 entries were run on the same input. Entries marked doc-verified were not run and are excluded from this table.

ToolVersion testedTranscript WERSummary factsTurnaroundSummary length
EchosawOur productEchosaw gamma pipeline, POST /v1/analyze/url19.3%34 sub / 146 del / 52 ins9 / 12all three financial figures correct178.8 ssingle analysis pass, transcript + summary + visual122 words
ScreenAppScreenApp web app, free plan; transcript provider "xai", summary provider "xai" as reported by the app20.2%48 sub / 134 del / 61 ins12 / 12six chapters + overview60.5 supload start to summary generatedAt507 words
OtterOtter web app, Basic plan, English (US) file import18.2%43 sub / 129 del / 46 ins6 / 12revenue target wrong (15M vs 50M)74 screated to modified timestamps207 words
Twelve Labsmarengo3.0 index (visual+audio) transcription + pegasus1.5 analyze, twelvelabs SDK 1.3.430.7%63 sub / 165 del / 140 ins (duplicated words)8 / 12no financial figures125.8 s3.5 upload + 100.8 index + 21.4 analyze99 words

How we evaluated

The video is the first 600 seconds of AMI meeting ES2002a (corner-camera view with the meeting-room audio track, H.264/AAC). The original 352x288 file was rejected by Twelve Labs as too low a resolution, so every service received the same 640x480 upscale (ffmpeg, libx264 CRF 20, audio stream copied). Each service was run once, on the test date, from a free or free-tier account created with an echosaw.com address; Echosaw ran on our own production API.

Transcript WER is computed with jiwer against the 1,200-word AMI reference transcript after both texts pass through the Whisper English normaliser (lowercasing, punctuation, number and contraction normalisation, "um"/"uh" dropped), identical to the speech-to-text list. Summary facts is the number of twelve facts stated in the meeting (purpose, brief, the four roles, the three-stage process, the whiteboard icebreaker, the 25 euro price, the 50 million euro target, the 12.50 euro cost cap, international considerations) that appear in the service's own summary, matched by pattern in score.py. It measures coverage, not writing quality.

Turnaround is wall-clock seconds from the start of upload to summary available. For Twelve Labs and Echosaw it is measured by the runner; for ScreenApp it is the noted upload start to the summary's generatedAt timestamp; for Otter it is the conversation's created to last-modified timestamps as reported by its API. ScreenApp and Otter have no public API on their free plans, so their transcript and summary were captured from the JSON their web apps load and fed to the runner with --import; the capture steps are documented in run.py.

Pricing is the list price on each vendor's public pricing page on the test date. Google Video Intelligence and Azure Video Indexer were not run (they require a billed GCP project or Azure subscription that we did not set up for this round); they are ranked on published capability and price only and do not appear in the results table.

Dataset

AMI ES2002a, first 600 s, corner camera (640x480 upscale) Kick-off meeting of a four-person design team: introductions, a whiteboard icebreaker, project finance, and international considerations. Reference transcript (1,200 words) built from the AMI word-level annotations by the speech-to-text harness's build_dataset.py. License: AMI Meeting Corpus CC BY 4.0.

Harness and raw output: /curated-lists/harness/video/

Ranking rubric

Transcript accuracy35%
Word error rate against the human reference on the meeting video.
Summary quality30%
Facts from the meeting correctly present in the summary, and absence of wrong facts.
Turnaround and workflow15%
Time from upload to summary, and whether the result is reachable by API.
Price20%
List price for one 10-minute video on the cheapest plan that supports the workflow.

The rankings

#1

Echosaw

Our productMeasured

Orange Sky Software · Best for teams that want transcript, summary, timeline, and visual analysis from one API call, then search across the results

Got all three financial figures right and was the only service to describe what was on screen (the room, the whiteboard), at the cost of the slowest turnaround.

Strengths

  • 19.3% transcript WER, within about a point of the best result on this audio.
  • Summary covered 9 of 12 facts and stated all three financial figures correctly; it also identified the room and the whiteboard activity from the video track, which no audio-only service can do.
  • Single POST /v1/analyze/url returns transcript, summary, timeline, key phrases and billing in one report; no separate index step.

Limitations

  • 178.8 s turnaround for a 10-minute video, the slowest measured: the pipeline runs visual analysis and moderation whether or not you need them.
  • Summary omitted the "original, trendy, user-friendly" brief and the three-stage process, and did not give the industrial designer's role.
  • $0.43/min audio+video on the Agency plan ($4.30 for this video) is the highest per-minute rate in the test; the Starter plan is $0.58/min. Audio-only analysis is $0.22-0.37/min.

Pricing: Subscription $9-49/month plus usage: audio+video $0.58/min (Starter) to $0.43/min (Agency); audio-only $0.37 to $0.22/min

#2

ScreenApp

Measured

ScreenApp Pty Ltd · Best for individuals who want a chaptered summary of a recording fast

Fastest turnaround and the most complete summary in the test (12 of 12 facts, six chapters), but the free plan caps you at three recordings and there is no API below the Business plan.

Strengths

  • 60.5 s from upload to summary, the fastest measured.
  • Summary hit all 12 facts, including the correct 50 million euro target, with chapters and speaker attribution.
  • Free plan needs no card and includes the full transcript; a 24.9 MB upload processed on the free tier.

Limitations

  • 20.2% transcript WER, mid-pack; the transcript rendered one participant's name two different ways.
  • Free plan allows 3 recordings total and 1 transcription a month; transcript download and export are paid features (we captured ours from the app's own JSON).
  • API access requires the $34/month Business plan.

Pricing: Free (3 recordings, no card); Growth $19/month billed annually ($228/year); Business $34/month billed annually with API access

#3

Otter

Measured

Otter.ai · Best for meeting notes from Zoom, Teams, and Meet where a bot joins the call; file import is secondary

Best transcript accuracy in the test and a fast turnaround, but the summary misstated the revenue target and the free plan allows only three file imports for the life of the account.

Strengths

  • 18.2% transcript WER, the lowest measured on this meeting audio.
  • 74 s from upload complete to processing done; summary, outline, and keyword list generated automatically.
  • Basic plan is free with 300 transcription minutes a month and no card.

Limitations

  • Summary stated the revenue goal as 15 million euro; the recording says fifty million. It also named participants without their roles, so it hit 6 of 12 facts.
  • Three lifetime file imports on Basic; Pro ($16.99/user/month) allows 10 imports a month.
  • No self-serve API; results were captured from the web app.

Pricing: Basic free (300 min/month, 3 lifetime imports); Pro $16.99/user/month or $8.49 billed annually; Business $30/user/month or $24 annually

#4

Twelve Labs

Measured

Twelve Labs · Best for developers building video search or open-ended video Q&A who want a prompt-driven summary rather than a fixed report

A developer API with a generous free tier and a prompt you control, but the transcript returned by the index carried duplicated words that doubled its error rate.

Strengths

  • Pegasus summary was accurate on everything it stated (roles, brief, icebreaker, international sales) and correctly noted no decisions were taken.
  • Free plan: 600 minutes of indexing, no card; indexes persist 90 days.
  • Summary prompt is yours; the same indexed video can then be searched with Marengo.

Limitations

  • 30.7% transcript WER: the transcription attached to the index repeated many words across adjacent segments ("this this is is the the kickoff kickoff"), producing 140 insertions. Speech transcription is a by-product of indexing, not a first-class output.
  • Summary left out all three financial figures and the three-stage process (8 of 12 facts).
  • Two-step workflow (index, then analyze); 125.8 s total of which 100.8 s was indexing.

Pricing: Free 600 minutes; Developer pay-as-you-go: Marengo indexing $0.042/min, Pegasus analyze $0.0292/min input + $0.0075 per 1k output tokens (about $0.72 for this video)

Microsoft Azure · Best for azure-committed organisations that need transcription, translation, and visual insights under one Azure resource, including edge deployment via Arc

The broadest insight catalogue of any service here (speaker indexing, OCR, faces, topics, generative summaries) and the only one deployable on-premises; not measured in this round.

Strengths

  • Transcription with speaker indexing, translation, topics, named entities, and generative textual summaries per the product documentation.
  • Runs in the cloud or on Arc-enabled Kubernetes for data-residency requirements.
  • Trial account: up to 10 hours of free indexing for website users and 40 hours for API users, per the pricing page.

Limitations

  • Not run in this test; WER, summary coverage, and turnaround figures are unavailable.
  • Pricing is per input minute by preset (Basic, Standard, Advanced audio and video); the rates are shown only after selecting a region on the pricing page, so no single list price is quoted here.
  • Requires an Azure subscription and a Video Indexer account bound to a storage account before the first upload.

Pricing: Per input minute by preset (Basic / Standard / Advanced audio and video); rates vary by region. Trial: 10 h free (web), 40 h free (API)

Google Cloud · Best for GCP pipelines that need labels, shots, and text detection alongside a transcript and will build their own summary step

A feature-per-minute annotation API with 1,000 free minutes a month of each feature; it returns a transcript but no summary, so it is a component rather than a complete service.

Strengths

  • Speech transcription $0.048/min after 1,000 free minutes per month per feature, per the pricing page.
  • Label, shot, explicit-content, object, text, logo, face, and person detection in the same request.

Limitations

  • Not run in this test; WER and turnaround figures are unavailable.
  • No summarisation feature: producing a summary requires a second call to a language model (e.g. Gemini on Vertex AI), which is priced separately.
  • Requires a GCP project with billing enabled even to use the free tier.

Pricing: Per feature per minute after 1,000 free minutes/month: speech transcription $0.048, label detection $0.10, shot detection $0.05, object tracking $0.15

Frequently asked questions

Which service produced the most accurate transcript?
On our 10-minute AMI meeting video, Otter had the lowest word error rate (18.2%), followed by Echosaw (19.3%) and ScreenApp (20.2%). Twelve Labs was 30.7% because the transcription attached to its index repeated words across segment boundaries. All four are within the range we measured for dedicated speech-to-text APIs on the same audio (16-33%).
Which service produced the best summary?
ScreenApp covered all 12 facts we checked, with chapters and speaker names. Echosaw covered 9 and was the only service to describe the room and whiteboard from the video track. Twelve Labs covered 8 with no errors but omitted every number. Otter covered 6 and misstated the revenue target as 15 million euro instead of 50 million, the only factual error in any summary.
Why is Echosaw ranked first when it was not the most accurate transcript?
The rubric weights transcript accuracy 35%, summary quality 30%, turnaround and workflow 15%, and price 20%. Echosaw was within 1.1 points of the best WER, got every number in the summary right, and is the only measured service with a self-serve API. It lost on turnaround (slowest) and price (highest per minute). ScreenApp had the better summary and speed but no API below its Business plan. Echosaw is our product; the raw outputs are published so you can weigh the rubric differently.
Otter alternative: which services did better on our meeting video?
Otter had the best transcript (18.2% WER) but the weakest summary (6 of 12 facts, with the revenue target misstated as 15 million euro instead of 50 million), and its free plan allows three file imports for life with no self-serve API. ScreenApp covered all 12 facts in 60.5 s at 20.2% WER, and Echosaw covered 9 of 12 with every financial figure correct at 19.3% WER and is the only measured service with a self-serve API. Twelve Labs had a worse transcript (30.7% WER) but a better summary (8 of 12 facts, no errors) and a free 600-minute API tier.
Why were Google and Azure not tested?
Google Video Intelligence needs a GCP project with billing enabled and Azure Video Indexer needs an Azure subscription; neither offers a card-free path. We did not set those up for this round, so both are ranked on documented capability and price and marked documentation-verified.
How can I reproduce these results?
The runner (run.py), scorer (score.py), and the raw transcript, summary, and timing JSON for each service are published at echosaw.com/curated-lists/harness/video/. Twelve Labs and Echosaw run from API keys you supply as environment variables; ScreenApp and Otter have no free-plan API, so run.py accepts the JSON their web apps load via --import. The video is the first 600 seconds of AMI ES2002a, upscaled to 640x480 with the ffmpeg command in the methodology.

Ready to bring powerful multimodal AI to your media operations?

Trusted at scale to extract semantic insights, build intelligent timelines, deliver accurate transcripts, analyze audio and visual content, and generate synthetic media — with full control and security. Start with our Starter plan for $9/month — usage-based pricing so you only pay for what you analyze.