Intelligent video analytics, layer by layer

Detection, transcription, on-screen text, and scoring in one report.

Echosaw combines scene and object detection, transcripts, on-screen text, and rubric scoring into timestamped findings.

Analyze footage free

Works with your own files — not just YouTube links.

Encrypted in transit (TLS 1.2+) and at rest (AES-256)Cited, timestamped answersFiles up to 5GB
Objects and scenesTranscriptScreen textTimestamped findings

What intelligent video analytics means in Echosaw

Echosaw is multimodal AI intelligent video analytics software: intelligent video analytics in Echosaw is the combination of object and scene detection, speech transcription, on-screen text reading, and a language model that scores footage against the events you define. Echosaw is one engine: it ingests the footage, analyzes it, scores it against the event rubric or policy you enter, and summarizes what it found.

What you upload

Recorded video in MP4, MOV, AVI, MKV, or WebM up to 5 GB per file, plus the events or criteria you want checked, written in plain language.

See It In Action

Real meeting. Real analysis. Real insights.

Explore a full Echosaw analysis of a recorded meeting: timestamped transcript, key moments, AI insights, generated outputs, and a chat grounded in the video.

  1. 1. Upload video, audio, images, or documents
  2. 2. Analyze transcripts, timelines, and key moments
  3. 3. Ask questions and get cited answers

What comes back

Findings marked Criterion Met, Criterion Unmet, or Flagged for Review at their timestamps, a Timeline of detected objects and scenes, a Transcript, a Screen Text tab, content warnings, and search across your library. Results are available in the web app and through the REST API and MCP server, so you can push findings into your CRM, ticketing, or VMS.

How Echosaw fits your workflow

Echosaw raises each event at its timestamp with the evidence attached; your team decides and acts. It reports what happened in the footage and when; face detection marks where people appear so your team can identify them from the evidence, and identity matching is available on request. Echosaw analyzes footage as it arrives — record in the browser, upload, or send clips through the API — and returns findings as close to real time as recorded analysis gets. Need continuous camera-feed ingestion? Contact us.

The layers of intelligent video analysis

LayerWhat it doesWhere you see it
Scene detectionFinds where the picture changes and picks keyframesKeyframes used for visual analysis and search
Object and scene detectionLabels what is in each part of the videoTimeline
On-screen textReads text visible in the frameScreen Text tab
SpeechTranscribes the audio track with timestampsTranscript tab
ModerationFlags harmful visual and spoken contentInsights tab and Timeline
ScoringChecks the footage against the events you defineCriterion Met, Unmet, or Flagged for Review, each with a timestamp

From detection to a decision

Detection alone produces a pile of labels. Intelligent video analytics turns those labels into an answer. In Echosaw, the detection layers describe the footage — objects, scenes, text, speech — and a scoring step reads that description against your criteria. The result is a list of findings you can act on, each tied to the second in the clip where it happens.

This is the same engine behind every Echosaw report: ingest, analyze, score against a rubric or policy, summarize. For surveillance footage the rubric is your list of events, and the evidence is the footage; for moderation it is a fixed harm policy.

Why evidence matters in intelligent video analysis

A finding is only useful if someone can check it. Click any finding and the clip plays from that moment, so the person reviewing it watches the few seconds that matter and confirms or rejects the call. That keeps a human in charge of the decision while the software does the watching.

Where it fits

Intelligent video analytics in Echosaw fits review work on footage as it arrives: incident review, checking a shift’s clips for specific events, auditing footage against a written procedure, or answering questions about a recording in Ask Emma. For the full surveillance picture, see AI video analytics for surveillance footage.

Intelligent video analytics FAQ

What is intelligent video analytics?
Intelligent video analytics (IVA) is software that understands what happens in video, not just whether pixels changed. Instead of a motion alert, it tells you what was in the scene, what was said, and whether an event you care about happened, and when.
How is intelligent video analysis different from motion detection?
Motion detection reports that something moved. Intelligent video analysis reports what it was and whether it matters: Echosaw detects objects and scenes, reads on-screen text, transcribes speech, and scores the footage against the events you define, each with a timestamp.
Does Echosaw run intelligent video analytics on live cameras?
Echosaw analyzes footage as it arrives — record in the browser, upload, or send clips through the API — and returns findings as close to real time as recorded analysis gets. Need continuous camera-feed ingestion? Contact us. Every finding carries its timestamp in the clip.
What does it cost?
Plans are $9, $19, $29, and $49 per month plus usage pricing. Agency, at $49 per month, has the lowest rates: $0.43 per minute for audio plus video and $0.22 per minute for audio only, for videos up to 210 minutes. Example: on Agency, a 10-minute clip costs 10 × $0.43 = $4.30 in usage on top of the $49 monthly plan. Your first three uploads are free.

Last updated: September 28, 2026

Ready to bring powerful multimodal AI to your media operations?

Trusted at scale to extract semantic insights, build intelligent timelines, deliver accurate transcripts, analyze audio and visual content, and generate synthetic media — with full control and security. Start with our Starter plan for $9/month — usage-based pricing so you only pay for what you analyze.