Update
September 9, 2026 · 9 min read
By Tony Cheng Tong
In April, before NAB Show, I wrote that Zonic was becoming an operational AI layer between raw footage and a professional production pipeline.
That description was directionally right, but it was incomplete.
Five months of building with editors, production workflows, long recordings, different devices, and increasingly capable AI systems have clarified the larger ambition: Zonic is a video-intelligence company, and video intelligence is an expert–AI interaction problem.
The goal is not to ask AI for a finished video and remove the expert from the process. The goal is to give an expert a system that can see, hear, retrieve, compare, reason, propose, revise, and operate on footage—while keeping the expert's judgment visible and in control.
That is what we have been building since April.

The direction: AI turns footage into evidence and options; the expert decides what the story becomes. This is an editorial illustration, not a product screenshot.
Video contains more than pictures. It contains speech, people, actions, places, chronology, emotion, camera decisions, repeated takes, technical quality, and the intention of whoever recorded it. Understanding a long recording—or hundreds of clips from one production—means combining all of those signals.
Editing adds another layer. Finding a relevant moment is not the same as knowing whether it belongs in a story. A fluent interview answer may be factually redundant. A visually imperfect shot may be the only honest record of an important event. Two cameras may show the same moment, but only an editor understands which angle preserves continuity and meaning.
This is why we do not believe professional video intelligence should be packaged as a one-click replacement for human craft.
Zonic is designed as an interaction between two forms of intelligence:
The interface between them is not only a chat box. It is the footage library, the transcript, the evidence behind an answer, the editable sequence, the rendered frame, the professional project file, and the record of what the expert changed.
The product has become clearer as two reusable intelligence systems.
Zonic Nandemo is the perception layer. It turns raw video, images, and audio into structured understanding: scenes, speech, speakers, people, labels, descriptions, key quotes, metadata, and searchable time ranges. Different models are useful for different signals, so Nandemo is built to coordinate capabilities instead of pretending that one model sees everything equally well.
Zonic Dekiru is the reasoning and action layer. It works across Nandemo's evidence, answers questions, retrieves the right moments, applies editorial skills, and builds an editable sequence. It can inspect composed frames, identify problems before committing a change, and revise a sequence incrementally instead of regenerating an entire result whenever the expert asks for a correction.
Together, they power two products. Zonic Workbench is the interactive environment for footage, conversation, review, sequencing, rendering, and export. Zonic Platform API exposes the same intelligence foundations to connected applications and headless workflows.
The product surface can expand. The intelligence does not need to fragment.
This was not one continuous feature launch. Each stage removed a different barrier between an AI demonstration and an operational product.
We separated scene detection, scene understanding, speech processing, embeddings, and face discovery so that each part could scale and fail explicitly. We added more provider choices, long-file upload handling, transcription-language controls, source retrieval, project cloning, metadata export, and the foundations of local execution.
The lesson was simple: an agent cannot be more reliable than its evidence. Before asking AI to edit, we had to make the source material observable, structured, and recoverable.
We introduced the Google Cloud execution path and migrated the editing agent into a Python-based runtime that could run as a cloud job. Long videos gained windowed analysis, compression and codec normalization became part of ingest, and transcript intelligence expanded to speaker diarization and key-quote extraction.
We also began treating development and production as separate operating environments. That sounds ordinary; becoming ordinary was the point. Video intelligence only matters if it can survive large files, slow networks, long-running work, provider failures, and a browser that disconnects.
The sequence became the center of the system. We added human-readable scene breakdowns, multitrack structure, audio alignment, subtitles, music, transcript correction, comments and review, source sharing, and more faithful Premiere Pro output. Cloud and Hybrid execution began converging around the same project while originals could remain close to the editor.
This changed Zonic from a system that could describe footage into one that could operate on it and return a result an editor could continue working with.
We extracted shared packages for the timeline, source management, provider capabilities, pricing, and the Dekiru agent. The agent gained long-running cloud jobs, interruption and recovery, visual inspection of the actual composed sequence, task-specific skill cards, and personal skills that a user can review and reuse.
The Workbench gained a flexible professional workspace, mobile layouts, shareable review surfaces, richer image and audio understanding, multilingual output controls, Google Drive workflows, API documentation, account connections, and usage metering. CapCut and JianYing joined Premiere as editing destinations.
These are not all equally visible in a demo. Together, they form the contracts that let more than one product depend on the same intelligence safely.
We unified the Cloud and Hybrid product model further, verified DaVinci Resolve export against a real Resolve installation, and brought captions and graphics into much closer agreement across preview and professional editing targets.
We also verified the first Canva ingest workflow in Canva's development environment: a Canva user can connect an existing Zonic account, select media from a design, send it into a Zonic project, and see the analysis complete. Writing an editable rough cut back into Canva is the next step, not a shipped claim today.
Alongside the Workbench, our EagleShot mobile work explores the other end of the workflow: video intelligence beginning at capture, including phone and AI-glasses experiences. Hardware-specific checks remain deliberately separate from what we claim as proven.
A recurring engineering principle now guides Zonic: one sequence should not become a different creative decision merely because it is viewed or exported somewhere else.
The same sequence can be inspected in the Workbench, rendered as a preview, or handed to a professional editor. Today that path includes Premiere Pro, CapCut and JianYing, and a locally verified DaVinci Resolve workflow. Subtitles, canvas formats, multiple video and audio lanes, graphics, transitions, and source transforms all make that promise harder than a list of export buttons suggests.
So we built a shared sequence compiler and explicit capability checks. Where editing applications interpret the same instruction differently, we test real output and report limitations instead of silently dropping creative intent.
This is part of the product's ambition. Expert–AI interaction must extend all the way to the tools experts trust.
Zonic now has several possible front doors, but one underlying thesis.
For professional teams, the immediate opportunity is footage-heavy work: productions, studios, archives, events, and broadcasters where finding and assembling the right material is a real operational bottleneck. We are exploring an operator-approved, near-live workflow with television teams—AI continuously understands incoming footage and proposes a rolling sequence while capture and final editorial control remain human-owned.
For individual creators, intelligence should appear inside the tools and devices they already use. Canva is one path. Mobile and AI-glasses capture are another. A creator may never want to manage a professional post-production workspace, but they still benefit from the same scene understanding, retrieval, reasoning, and editorial skills.
For enterprises with sensitive media, Hybrid execution already keeps full-resolution work close to local files while scalable intelligence uses the cloud. A fully offline, on-premises product is a direction we are designing; it is not something we claim to have finished.
These paths share Nandemo, Dekiru, the source model, the sequence contract, and the skill system. Each new interaction can improve the same foundation instead of creating an isolated application.
The long-term advantage is not a prompt that generates a timeline once. It is the structured relationship between footage, the AI's proposal, the expert's correction, and the final accepted work.
Zonic already supports explicit training workflows and personal editorial skills that users can inspect. Human review produces valuable correction signals today. The larger learning loop—decoding arbitrary finished professional projects and learning automatically from the difference—is still ahead of us.
We are careful about that distinction because trust matters. Editorial intelligence should not absorb one person's preference as a universal rule, learn from private work without consent, or hide why a decision changed. The future system must preserve attribution, provenance, review, and the ability to forget.
If we build that correctly, every expert–AI interaction can make the next collaboration more precise.
Our next stage is focused on turning the breadth we have built into repeatable outcomes:
We are beginning conversations with footage-heavy teams who want to design these workflows with us, distribution partners who can bring video intelligence into existing creative ecosystems, senior builders who want to work on expert-centered AI, and investors who share the ambition of building an intelligence layer for the world's video.
If that describes you, talk with us. We would like to understand where your experts lose time today—and what they could create if every frame were already understood.