ChatCut — a conversational AI video editor where every instruction becomes a normal edit on a real multi-track timeline

Describe the cut in plain language and ChatCut applies it to a real multi-track timeline — transcript editing, captions, and a free tier.

  • AI Video Editor
  • Conversational Editing
  • Transcript Editing
  • Talking-Head Video
  • Auto Captions
Publisher
ChatCut Inc.
Type
AI Video Editor
Pricing
Freemium
Reviewed
24 September 2026
Official site

Quick verdict

Use when

  • Your footage is talking heads, interviews, podcasts or screen recordings, so the cut is decided by what was said rather than by what is on screen
  • You would rather delete a sentence by deleting its words in the transcript than hunt for the right frames on a waveform
  • You want the agent to make the first pass and still pick up the same timeline by hand afterwards, with nothing locked behind the prompt
  • You want captions, a voiceover, music, motion graphics or a generated clip created inside the project instead of round-tripped through other tools
  • You already drive a coding agent and want it to operate a video editor rather than only write text
  • You need a rough cut of a long recording today and would rather describe the edit than learn shortcuts for it

Skip when

  • You need frame-accurate control of every cut, keyframe and speed ramp, where a keyboard is faster than a sentence
  • The edit is effects-led or motion-led, with timing judged by eye rather than by the spoken word
  • You need professional colour grading, multicam switching or finishing tools that a browser editor does not carry
  • Your footage cannot leave your own machines at all, or you have to work offline
  • You need an ongoing free plan rather than a starting balance that does not renew
  • The output is a single trim or resize, which a one-purpose utility handles without an account

Try instead

If you want a different flavour of AI editor, or the whole job is one small cut rather than an edit, those live here too.

ChatCut vs Descript vs CapCut vs VEED

The dividing line in AI video editing is what the model is allowed to touch: most tools render a template you then live with, while ChatCut turns instructions into ordinary clips on an editable timeline. Against Descript, which edits through a transcript, CapCut, which competes on template volume and platform-native effects, and VEED, which is a browser editor with AI features attached, the four differ on tiers, learning cost, output limits and where your footage is processed.

Tap a dimension to focus

Pricing

Roughly even
  • ChatCutThis page
    • A free tier with a one-time starting credit balance rather than a monthly renewal
    • The same free tier carries a cumulative cloud-export quota for the life of the account, not per month
    • Pro is a subscription that adds a recurring monthly credit allocation and removes the export quota
    • Generation consumes credits — video clips, images, music, voice and motion graphics are each metered — so a tier is a budget for output rather than a feature gate
    • Higher tiers scale the same recurring allowance rather than unlocking different features
    • A free tier with a small monthly allowance of media minutes and AI credits
    • Paid tiers are priced per seat and raise media hours, AI credits and export resolution together, so capacity is the main axis
    • Custom voice clones, avatar generation and timeline export to professional editors sit above the entry tier
    • Enterprise is quoted rather than listed, with custom credits, retention and legal terms
    • The core editor is free across web, desktop and mobile, with a Pro subscription selling effects, assets and export options
    • Pro is a flat subscription rather than a metered credit pack, so ordinary editing does not consume a balance
    • Some AI features are limited or metered separately from the subscription
    • Commercial licensing and team seats are handled as separate offerings from the consumer plan
    • A free tier that caps export length and marks the output
    • Paid tiers are per seat and scale export length, resolution and access to the wider AI toolset
    • Team plans add shared workspaces, brand controls and review tooling
    • Annual billing is cheaper than month-to-month on the same tiers, which changes the effective per-seat cost
  1. ChatCut — official site
  2. ChatCut — pricing and credits
  3. ChatCut — what the product is
  4. ChatCut — features
  5. ChatCut — web, desktop, and agent plugin
  6. ChatCut — free and Pro plans
  7. ChatCut — terms of service
  8. Descript — official site
  9. CapCut — official site
  10. VEED — official site

Preview of ChatCut - not the live app. Confirm details on the official site.

Does this overview help you decide?

—

(—)

Learn more

Details below the decision summary—features, workflow, and scope notes.

What is ChatCut?

Describe the cut in plain language and ChatCut applies it to a real multi-track timeline — transcript editing, captions, and a free tier.

What it costs

A free tier with a one-time credit balance and a cloud-export quota, then a Pro subscription that adds recurring credits and removes the quota.
Free tier
Yes
Pricing summary
The structure is a free tier plus a subscription, and the free tier has a shape worth understanding before you plan around it. It includes the editor and a one-time starting balance of credits rather than a monthly allowance, and the cloud-export side carries a cumulative quota across the life of the account instead of resetting each month. That combination is generous enough to evaluate the product properly and to finish a short project, but it is a starting position rather than a place to publish from indefinitely — once the balance and the quota are spent, the editor still opens but the work runs into a wall. Pro changes two things: it adds a recurring monthly credit allocation and it removes the export quota, while leaving the editor itself identical. That matters because the free tier is not a crippled product — the manual timeline, uploads and transcription are all included — so what you are buying is throughput rather than access. Credits are consumed by generation, and the things that consume them are the ones you would expect: video clips from models, images, music, voiceover and motion graphics. Because different models cost different amounts per unit of output, an allowance is better read as a budget for finished work than as a count of videos, and a project heavy on generated clips will burn through it far faster than one that is mostly cuts and captions. Note also that supported exports render locally in the desktop app while the web editor renders through the cloud workflow, so the export quota applies to the browser path; anyone doing volume work has a reason to install the desktop build for that reason alone. Tiers and credit allowances in this category are restructured often, so confirm the current plans and what each generation costs on the official pricing page rather than working from a number quoted elsewhere.

Reviewed on 24 September 2026 · ChatCut — pricing and credits

What ChatCut includes

The capabilities as the product documents them, read against what each one is actually for in an editing session.
  • Conversational editing on a real timeline

    The agent is a layer over a multi-track editor rather than a render button, so an instruction produces ordinary clips you can trim, move or delete by hand. That is the product’s central claim and the thing that separates it from template-based generators: the edit stays yours after the model has touched it. It is also the feature that sets the learning curve, because the useful skill is learning to describe an edit precisely enough that the result needs correcting rather than replacing.

  • Transcript editing, with the audio following the text

    Speech is transcribed and mapped word by word back to the timeline, so deleting a word in the transcript removes it from the audio and reordering paragraphs reorders the cut. For interview and podcast footage this is usually the fastest route to a first pass, and it collapses a class of edit — removing filler, tightening an answer, cutting a whole tangent — into something closer to editing a document than operating a waveform.

  • Media generated inside the project

    Video clips, images, motion graphics, music, sound effects and voiceover are all generated in place rather than imported from separate services. The practical effect is that a gap in the edit which would otherwise mean opening another tab becomes a prompt. Motion graphics matter more than they sound, because titles, lower thirds and logo animations are the kind of work that stalls a solo editor, and transparent ProRes export for those graphics is what lets you composite them over footage elsewhere.

  • Captions and narration as part of the cut

    Captions are generated with word-level timestamps across a wide language set and styled with presets, and narration is available from curated voices in 32-plus languages with direct placement on the timeline. Word-level timing is the detail that matters, because it is what keeps burned-in captions honest during fast speech, and having narration land on the timeline rather than in a separate file means revisions do not require a second export and import cycle.

  • Three product forms sharing one project

    The browser app, the native macOS and Windows editor, and the agent plugin that lets an external coding agent drive the editor all open the same cloud project. The differences are practical rather than conceptual: the desktop build can read local files and renders supported exports on your own machine, which sidesteps the cloud export quota, while the browser build is the fastest way to start and is the one to share. The plugin exists for people who already work inside a coding agent and would rather not leave it.

  • Exports that assume another tool downstream

    Beyond video, audio and subtitles, the editor exports XML for hand-off, which is the quiet acknowledgement that an AI-edited cut is often not the final one. It means the product can sit at the front of a pipeline ending in a professional editor rather than trying to replace it, which is a more honest position than products that present the generated render as a finished deliverable.

  • Multiple timelines per project

    A project can hold several timelines, so short versions, alternate cuts and platform-specific edits can be built from the same media without duplicating the source material. It is the structural feature behind repurposing one recording into several outputs, and it is why the product suits people whose material is a long recording rather than a single clip.

How to get a usable cut out of a conversational editor

The loop that works, and the four places it produces work instead of saving it.
  1. Start from the transcript, not the prompt

    The transcript is the fastest surface for the first pass on any spoken footage, because deleting a sentence there removes the sentence from the audio without you having to find its edges. Read through, cut what does not earn its place, and only then reach for the agent for the structural decisions — tightening an intro, reordering a section, building a short version. Using the agent for cuts you could make by deleting a paragraph is where the time goes.

  2. Describe the edit the way you would brief a person

    An instruction that would be ambiguous to a colleague will be ambiguous here, and the result is not a rough cut but a mess to undo. Say what the edit is for and what should not survive — the length, the audience, the parts to protect — rather than naming a single action and hoping the intent travels with it. Precision is the skill this tool actually demands, and it is worth spending the first session learning what specificity it needs.

  3. Keep the manual edit as the finish, not the fallback

    The value of a real timeline is that the agent’s output is a starting point you refine. Treat the generated cut as a first assembly and do the last pass by hand: fixing a hard cut on a breath, trimming a beat that reads as hesitation, checking that a caption lands where the emphasis is. A workflow that expects the prompt to be final will be disappointed; one that uses it to skip the boring 70 per cent will not be.

  4. Verify captions and pronunciation before publishing

    The failures audiences actually notice in AI-assisted edits are textual and auditory — a name mispronounced by a generated voice, a caption that drifts during fast speech, an emphasis landing on the wrong word, transcription confidently wrong on a proper noun or a term of art. These are cheap to catch on a read-through and expensive to miss on a published video, and no model quality removes the need for that check.

  5. Match the surface to the work

    Use the browser when starting or sharing, because it needs no install and the project is in the cloud either way. Use the desktop app when the media is local or the export is heavy, since supported exports render locally and that avoids both the upload and the cloud export quota. Use the agent plugin if you already work inside a coding agent. All three open the same project, so this is a workflow choice rather than a commitment, and there is no penalty for mixing them.

  6. Check what a generation will cost before building around it

    Credits are consumed by generation, and different models cost different amounts for the same duration of output. A project that leans on generated clips is a different purchase from one assembled from existing footage with captions added, and that difference is easy to discover halfway through a deadline. Plan the generated parts deliberately — where the footage genuinely needs creating rather than cutting — and spend the allowance on the shots that earn it.

Who ChatCut is for

The editing jobs this product actually maps onto.
  • Interview and podcast editors on a deadline

    This is the clearest fit, because the decisions in a conversation-driven edit are verbal and the transcript turns them into text operations. Removing filler, tightening an answer, cutting a tangent and building a shorter version from one recording are all faster through the transcript than through a timeline. The constraint to know is scale: the product is built for individual cuts and repurposing rather than for an hour-long production with multicam and finishing.

  • Small teams repurposing long recordings

    One recording becoming a full episode plus several short versions is the workflow the project structure is designed around, with multiple timelines in one project and captions generated in place. For a team of one producing across several platforms, that shortens the distance between having the material and having something to post, which is where small content operations usually stall.

  • People who make tutorial and screen-recording content

    Screen recordings share the property that makes transcript editing work: there is narration, and the narration is where the structure lives. Dead air, retakes and wrong turns are all easy to see in text and tedious to find on a waveform, and captions generated with word-level timing are the same asset that makes a tutorial watchable with the sound off.

  • Teams editing in a browser by default

    A cloud project that opens in a browser with no install is the least friction a collaborator can be given, and the desktop app and agent plugin open the same project for the people who want them. For a team that reviews or hands off cuts across machines, that consistency removes a class of file-juggling that desktop-only editors impose.

  • Developers who want an agent to operate the editor

    The agent plugin exists so an external coding agent can drive the editor and keep the result editable in ChatCut, which is a specific and unusual capability rather than a marketing line. Anyone already building workflows around a coding agent can close the loop on video, while still ending up with a timeline a person can finish rather than a render nobody can change.

When ChatCut is the right pick

The promise of AI video editing has mostly been delivered as a render button: you describe something, a template-shaped video comes back, and your influence ends there. That is genuinely useful when the job is producing something disposable at volume, and it is the wrong shape when the job is an actual edit, because the parts you want to change are precisely the parts the model decided. ChatCut takes the other position, and it is the interesting thing about the product: the agent does not replace the editor, it operates one. An instruction becomes clips on real tracks, which means the output of a prompt is a rough cut you can argue with rather than a file you have to accept. For footage where the decisions are verbal — interviews, podcasts, talking heads, screen recordings — that is a good fit, because the transcript gives you a second handle on the same timeline and deleting a sentence becomes a one-word edit. The trade is real and worth stating plainly. Conversational editing is slower than a keyboard once you know what you are doing, so this is a tool for the early passes of an edit rather than for fine control. Effects-led and motion-led work does not suit it. Anything that cannot leave your machines is out entirely, since the project lives in the cloud. And the free tier is a starting position rather than a permanent one, which rules it out if you need a free plan to publish from. Where it earns its place is the middle of the workflow that nothing else covers well: a long recording, a deadline, and a need for a watchable cut before the interesting decisions get made. It is worth noting too that the three forms — desktop, browser and the agent plugin that lets an external coding agent drive the editor — are one product with one project, so a team differently comfortable with each is not a problem. If the constraint is frame-accurate finishing, stay with a traditional editor. If it is turning a recording into a cut today, this is one of the few tools in the category built for exactly that.

Platform and product notes

What it is, what it will not do, and the facts worth verifying at the source.
An editor with an agent, not a generator with an export button
The distinction is structural rather than stylistic: instructions become clips on real tracks, so the output is editable and the product sits alongside editors rather than in place of them. That is what makes it useful for an actual edit and less useful for someone who wants the final render decided for them. If you want a finished video from a prompt with no timeline involved, a template-based generator is the closer match.
Cloud projects, with a local render path
Projects are stored in the cloud, which is what makes the browser, desktop and agent-plugin forms interchangeable, and it also means media and transcripts rest on the vendor’s side. Supported exports render locally in the desktop app, which is a meaningful distinction for anyone counting exports, but it is not an offline or on-device product: the project is remote and generation runs through third-party models.
Credits are metered by generation, and model prices move
Video, image, music, motion-graphics and voice generation each consume credits, and different underlying models cost different amounts for the same unit of output. The consequence is that the cost of a project is not fixed by its length — a clip-heavy edit and a caption-heavy edit of the same duration are different purchases, and a plan can become more or less useful as the models behind it change. Budget in credits against your own typical project rather than against a generic figure.
The free tier is a starting balance, not a monthly allowance
Free includes the editor and a one-time credit balance, and the cloud-export side is capped cumulatively over the account rather than resetting each month. That is enough to judge the product on real work, which is the point, but it does not support publishing indefinitely without paying. Anyone whose requirement is a free plan in perpetuity should treat that as a disqualifier rather than an inconvenience.
Rights and permissions for uploaded media are yours
The terms place responsibility for having the necessary rights, copyrights and permissions on whatever you upload or export with the account holder, and the terms of use state that user media is not used to train AI models. The second is the specific commitment worth reading at the source, since it is the one that changes what the product is acceptable for in a work context. Access is intended for users who are at least 18.
What it does not do
It is not a finishing tool: there is no professional colour grading, no multicam switching, and no substitute for frame-accurate control when an edit turns on a few frames. It does not work offline, and it is not for footage that cannot leave your machines. It does not currently offer a permanent free plan. And it does not make the editorial decisions — the transcript and the timeline still expect someone to decide what the cut should be.

Frequently Asked Questions

Quick answers about this tool—open a question to read more.