Intelligent video analytics, layer by layer
Detection, transcription, on-screen text, and scoring in one report.
Echosaw combines scene and object detection, transcripts, on-screen text, and rubric scoring into timestamped findings.
Works with your own files — not just YouTube links.
What intelligent video analytics means in Echosaw
Echosaw is multimodal AI intelligent video analytics software: intelligent video analytics in Echosaw is the combination of object and scene detection, speech transcription, on-screen text reading, and a language model that scores footage against the events you define. Echosaw is one engine: it ingests the footage, analyzes it, scores it against the event rubric or policy you enter, and summarizes what it found.
What you upload
Recorded video in MP4, MOV, AVI, MKV, or WebM up to 5 GB per file, plus the events or criteria you want checked, written in plain language.
Real meeting. Real analysis. Real insights.
Explore a full Echosaw analysis of a recorded meeting: timestamped transcript, key moments, AI insights, generated outputs, and a chat grounded in the video.
- 1. Upload video, audio, images, or documents
- 2. Analyze transcripts, timelines, and key moments
- 3. Ask questions and get cited answers
What comes back
Findings marked Criterion Met, Criterion Unmet, or Flagged for Review at their timestamps, a Timeline of detected objects and scenes, a Transcript, a Screen Text tab, content warnings, and search across your library. Results are available in the web app and through the REST API and MCP server, so you can push findings into your CRM, ticketing, or VMS.
How Echosaw fits your workflow
Echosaw raises each event at its timestamp with the evidence attached; your team decides and acts. It reports what happened in the footage and when; face detection marks where people appear so your team can identify them from the evidence, and identity matching is available on request. Echosaw analyzes footage as it arrives — record in the browser, upload, or send clips through the API — and returns findings as close to real time as recorded analysis gets. Need continuous camera-feed ingestion? Contact us.
The layers of intelligent video analysis
| Layer | What it does | Where you see it |
|---|---|---|
| Scene detection | Finds where the picture changes and picks keyframes | Keyframes used for visual analysis and search |
| Object and scene detection | Labels what is in each part of the video | Timeline |
| On-screen text | Reads text visible in the frame | Screen Text tab |
| Speech | Transcribes the audio track with timestamps | Transcript tab |
| Moderation | Flags harmful visual and spoken content | Insights tab and Timeline |
| Scoring | Checks the footage against the events you define | Criterion Met, Unmet, or Flagged for Review, each with a timestamp |
From detection to a decision
Detection alone produces a pile of labels. Intelligent video analytics turns those labels into an answer. In Echosaw, the detection layers describe the footage — objects, scenes, text, speech — and a scoring step reads that description against your criteria. The result is a list of findings you can act on, each tied to the second in the clip where it happens.
This is the same engine behind every Echosaw report: ingest, analyze, score against a rubric or policy, summarize. For surveillance footage the rubric is your list of events, and the evidence is the footage; for moderation it is a fixed harm policy.
Why evidence matters in intelligent video analysis
A finding is only useful if someone can check it. Click any finding and the clip plays from that moment, so the person reviewing it watches the few seconds that matter and confirms or rejects the call. That keeps a human in charge of the decision while the software does the watching.
Where it fits
Intelligent video analytics in Echosaw fits review work on footage as it arrives: incident review, checking a shift’s clips for specific events, auditing footage against a written procedure, or answering questions about a recording in Ask Emma. For the full surveillance picture, see AI video analytics for surveillance footage.
Intelligent video analytics FAQ
- What is intelligent video analytics?
- Intelligent video analytics (IVA) is software that understands what happens in video, not just whether pixels changed. Instead of a motion alert, it tells you what was in the scene, what was said, and whether an event you care about happened, and when.
- How is intelligent video analysis different from motion detection?
- Motion detection reports that something moved. Intelligent video analysis reports what it was and whether it matters: Echosaw detects objects and scenes, reads on-screen text, transcribes speech, and scores the footage against the events you define, each with a timestamp.
- Does Echosaw run intelligent video analytics on live cameras?
- Echosaw analyzes footage as it arrives — record in the browser, upload, or send clips through the API — and returns findings as close to real time as recorded analysis gets. Need continuous camera-feed ingestion? Contact us. Every finding carries its timestamp in the clip.
- What does it cost?
- Plans are $9, $19, $29, and $49 per month plus usage pricing. Agency, at $49 per month, has the lowest rates: $0.43 per minute for audio plus video and $0.22 per minute for audio only, for videos up to 210 minutes. Example: on Agency, a 10-minute clip costs 10 × $0.43 = $4.30 in usage on top of the $49 monthly plan. Your first three uploads are free.
Last updated: September 28, 2026
Ready to bring powerful multimodal AI to your media operations?
Trusted at scale to extract semantic insights, build intelligent timelines, deliver accurate transcripts, analyze audio and visual content, and generate synthetic media — with full control and security. Start with our Starter plan for $9/month — usage-based pricing so you only pay for what you analyze.