Diffusion Studio — an open-source AI video editor where a coding agent cuts footage through a CLI, on a canvas beside a non-linear timeline
Edit video on an infinite canvas and let a coding agent cut it through the CLI — free editor, unlimited 4K exports, no watermark.
- AI Video Editor
- Open Source
- Agentic Editing
- Motion Graphics
- Infinite Canvas
- Publisher
- Diffusion Studio Inc.
- Type
- AI Video Editor
- Pricing
- Freemium
- Reviewed
- 24 September 2026
Quick verdict
Use when
- You want the edit to exist as code, so a composition can be versioned, reviewed or re-run rather than rebuilt by hand
- You already work inside a coding agent and want it to open the project, inspect footage, cut, transcribe and render instead of only describing the edit
- The work is motion graphics, titles, explainers or scene-led pieces rather than a straight cut of long footage
- You need scenes that go beyond editing primitives — HTML, WebGL, 3D or shaders composited into the video
- You want a design-tool canvas and a non-linear timeline in the same project rather than choosing between the two
- You care that the editor is open source and that exports stay free and unwatermarked at any volume
Skip when
- You are not a developer and do not want a CLI in the loop, so the agent angle adds nothing you will use
- Your machine is not an Apple silicon Mac and you want the agentic workflows now, because those currently run through the desktop app
- You want a template-led consumer editor with a mobile app and preset-driven output
- You need collaborative review, shared brand templates or an approval workflow
- Your footage cannot be uploaded or routed to third-party model providers, since hosted AI runs in the cloud
- The job is a single trim, crop or format change, which a one-purpose utility handles without an account
Try instead
If you want a canvas-based creative tool, conversational editing without a CLI, or the whole job is one small operation on a file, those live here too.
Diffusion Studio vs Remotion vs CapCut vs VEED
Tap a dimension to focus
Pricing
Roughly even- Diffusion StudioThis page
- The editor, the CLI and the agent skills are free and open source on every plan
- Exports never consume credits, so unlimited output up to 4K costs nothing regardless of tier
- Credits are spent only when an AI model runs — generating media, analysing footage, transcribing
- The free tier carries a one-time starting credit balance rather than a monthly allowance
- Paid tiers scale a recurring monthly credit allowance, with volume discounts and optional top-ups, and a Teams tier pools credits with invoicing
- Free to use for individuals and organisations of up to three people
- A company licence becomes mandatory once the organisation reaches four people, which is the step most teams are surprised by
- Automation work is billed per render with a monthly minimum rather than per seat
- Manual creation seats are priced separately from automation, so a team can end up on two line items
- An enterprise tier adds custom terms, support and consulting
- The core editor is free across web, desktop and mobile, with a Pro subscription selling effects, assets and export options
- Pro is a flat subscription rather than a metered credit balance, so ordinary editing does not consume anything
- Some AI features are limited or metered separately from the subscription
- Commercial licensing and team seats are handled as separate offerings from the consumer plan
- A free tier that caps export length and marks the output
- Paid tiers are per seat and scale export length, resolution and access to the wider AI toolset
- Team plans add shared workspaces, brand controls and review tooling
- Annual billing is cheaper than month-to-month on the same tiers, which changes the effective per-seat cost
- Diffusion Studio — official site
- Diffusion Studio — pricing and credits
- Diffusion Studio — documentation
- Diffusion Studio — download for macOS
- Diffusion Studio — privacy policy
- Diffusion Studio — source repository (MPL 2.0)
- Diffusion Studio — terms and conditions
- Remotion — official site
- Remotion — licence and pricing
- CapCut — official site
- VEED — official site
Preview of Diffusion Studio - not the live app. Confirm details on the official site.
Does this overview help you decide?
—
(—)
Learn more
Details below the decision summary—features, workflow, and scope notes.
What is Diffusion Studio?
What it costs
- Free tier
- Yes
- Pricing summary
- This is one of the few products in the category where the free tier is not a trial of the product but the product itself. The editor, the CLI and the agent skills are free and open source, exports are unlimited up to 4K with no watermark, and the vendor states plainly that exporting never requires credits. Commercial use is permitted on every plan. What is metered is the AI: credits are spent when a model runs, which covers generating video, images, music, speech and sound effects, and analysing or transcribing footage so an agent can work from what is actually in it. That distinction is the one worth internalising, because it makes cost a function of how much you lean on AI rather than how much you edit — the vendor notes that building motion graphics usually costs nothing while editing footage usually means watching or transcribing it first. The free tier carries a one-time starting balance rather than a monthly allowance, so heavy AI use on the free plan is a limited resource. Paid tiers scale a recurring monthly credit allowance, with volume discounts at higher bundles and optional top-ups, and unused monthly credits do not roll over. There is also a refund position worth knowing before subscribing from the EU, where the statutory 14-day withdrawal right ends once the credits in a billing cycle have been consumed. None of this is quoted in figures here because the credit tiers and per-action costs are the parts that move, and the vendor publishes both in the app and on the pricing page, which is where to check what a specific generation will cost before committing.
Reviewed on 24 September 2026 · Diffusion Studio — pricing and credits
What Diffusion Studio includes
A canvas and a timeline in the same project
The workspace is an infinite canvas with a design-tool interface sitting beside a non-linear timeline that supports effectively unbounded layers and clips. The vendor gives a useful rule of thumb: use the canvas for generating and arranging assets in parallel, and the timeline for composing them with an agent that can see the current state of the composition. That split is why the tool does not feel like an NLE with a chat box bolted on — the two surfaces are for genuinely different kinds of work.
Edits that live in code
Compositions are declared in JSX and compiled into the project, and the vendor’s own summary of the product is that every edit stays in code. That is the claim that separates it from every timeline editor in this category, because it makes an edit something you can version, review in a diff, parameterise and re-run rather than something that exists only as a sequence of mouse actions. It is also the source of the learning cost: the payoff assumes a developer workflow.
An agent that operates the editor through a CLI
A command-line tool called dapi is the interface an agent uses, and skills teach a coding agent how to work with it. The documented surface is substantial — probing media, generating filmstrips and waveforms, transcribing and listening to footage, querying the scene graph, patching nodes, adding assets, and rendering — so the agent is editing a real project rather than producing a script for a human to follow. Support is stated for the major coding agents, and the vendor recommends letting an agent drive the desktop app directly rather than going through a plugin.
Scenes that are not limited to video primitives
Because scenes are built from web technology, compositions can use HTML, WebGL, 3D and shaders rather than only the effects an editor ships with. For explainers, title sequences and generative pieces this removes a ceiling that timeline tools hit early, and it is the capability most likely to be the actual reason to choose this over a conventional editor. Bezier curves with advanced easing cover the motion-graphics fundamentals on top of that.
Footage understanding as a first-class feature
An agent can watch footage and work from what is in it or what is said in it, which is what makes automated passes like finding the best moments, removing dead air and cutting filler words feasible. Scene search, summaries and quotes with timestamps come from the same capability. The practical consequence is that the model is being asked to understand the material rather than generate a substitute for it, and that is a more defensible use of AI in editing than prompt-to-video.
Generation, transcription and captions inside the project
Video, images, music, speech and sound effects can be generated from a prompt, and subtitles can be auto-generated in any language with timed, styled and animated text. Captions are the feature that quietly matters most for social output, and having generation and transcription in the same project as the edit means a revision does not require a second round trip through another tool.
Free, unwatermarked 4K exports
Unlimited exports up to 4K with no watermark are included on every plan, and exporting never consumes credits. No proxy step is needed for 4K playback. Combined with commercial use being permitted on the free plan and the editor being free rather than trialable, this is the most generous export position in the category by a clear margin, and it holds regardless of what you spend on AI.
An open-source editor you can inspect and build on
The web editor, desktop app, CLI and skills are published under MPL 2.0, so the client can be audited or run from source. For teams with procurement or security review, being able to read the code that touches their footage is a different proposition from taking a vendor’s word for it, and it is the reason the privacy comparison on this page reads the way it does.
How to run an agent-driven edit without losing the plot
Install the desktop app and connect the agent before you need it
Agentic workflows run through the macOS app on Apple silicon, and connecting an agent means installing the CLI and the skills once, through the app rather than by hand. The app has to stay open while the CLI works with it, because they talk over a local socket. Doing this deliberately at setup is worth it: the browser editor is a good editor, but it is not the product the agentic capabilities describe, and discovering that mid-project is the most common way people end up disappointed.
Let it watch and transcribe before it cuts
The capability that makes automated editing work is footage understanding, so the first useful thing an agent does is describe what is in the material rather than start cutting. Probe the media, generate filmstrips and waveforms, transcribe and search scenes, and read what comes back. Cuts proposed after that are grounded in the actual footage; cuts proposed before it are guesses dressed as edits, and they cost credits either way.
Split generation from composition
Use the canvas for producing and organising assets in parallel and the timeline for composing them, which is the division the documentation recommends and the one that keeps a project comprehensible. Generating assets is the part that consumes credits, so it is also where the cost concentrates; keeping it in one place rather than scattered through the timeline makes it possible to see what a project actually costs.
Treat the composition as code you will maintain
The whole argument for this tool is that an edit is a program, so it is worth getting the benefits of that deliberately: name the composition nodes sanely, keep parameters that vary as inputs rather than hard-coded values, and keep reusable pieces as skills. A composition built that way can be re-run against new footage or re-branded for a client in a way a timeline project cannot. One built by letting an agent patch nodes ad hoc ends up as unmaintainable as a hand-edited timeline, with less visibility.
Review transcriber output before trusting it
Transcription and footage analysis are the highest-value AI features here and also the ones whose errors are invisible until a published video. A mis-transcribed technical term becomes a cut in the wrong place, a scene labelled incorrectly moves the wrong clip into a highlight, and a caption drifts during fast speech. Read the output against the footage on anything that ships, and treat the transcript as an editing handle rather than as a record of what was said.
Plan AI spend against the free/metered split
Exports are free and unwatermarked at any volume, and the editor itself is free, so the only cost is AI models. That makes budgeting unusually legible — you are buying analysis and generation, not access — but it also means an agentic loop that transcribes and re-analyses repeatedly is spending real credit for a result a human could have reached by reading the transcript once. On the free tier the balance is one-time, so decide early which parts of the workflow must be AI and which are just faster when a model does them.
Who Diffusion Studio is for
Developers and creative engineers building video pipelines
This is the clearest fit and the reason the product exists: video as something a program produces, with compositions in JSX, a CLI an agent can drive, and reusable skills. Parameterised explainers, data-driven spots and template systems that regenerate against new inputs are all natural here, and none of them are achievable in a timeline editor without rebuilding the edit every time.
Motion designers who already write code
The canvas, bezier curves with advanced easing, and scenes built from HTML, WebGL, 3D and shaders cover motion-graphics work that conventional editors handle badly or not at all. For a designer comfortable in code, the combination of a visual canvas for arranging and compositions for precision is a better fit than either a pure code framework or a pure timeline.
Teams that need video output they can reproduce
When the same video has to be produced repeatedly with different data, branding or locale, an edit stored as code is the only version that scales. Being able to re-run a composition rather than redo it is the difference between a template that works and a template someone maintains by hand, and it is where the licence-free editor and free exports compound.
Anyone whose budget rules out per-seat editors
The editor is free, exports are free and unwatermarked at 4K, and commercial use is permitted on every plan, with AI credits the only metered cost. For a small studio, an educator or a side project, that combination removes the subscription decision entirely and leaves only the question of how much generative AI is worth paying for.
Agent-first workflows where a human finishes the cut
An agent can open a folder of footage as a project, understand it, and assemble a first pass that a person then refines on the timeline, with both working on the same project. Teams already running coding agents on other parts of their pipeline can add video as another step rather than another tool to learn.
When Diffusion Studio is the right pick
Platform and product notes
- The client is open source; the hosted service is not
- The web editor, desktop app, CLI and skills are published under MPL 2.0, so the software can be inspected, contributed to or built from source. The server side is a hosted commercial service, and the open-source licence covers the software rather than the videos you make with it. That distinction matters when weighing the privacy position: you can read the client, but using the hosted editor still means projects and generated assets live in the vendor’s cloud.
- Agentic workflows currently require the macOS desktop app
- The browser editor runs on Windows, macOS and Linux with nothing to install, and that is the version most people should try first. Coding-agent support and agentic editing workflows, however, currently run through the desktop application, which requires macOS 12 or later on an Apple M1 chip or newer. A Windows desktop app is stated as planned but not available. Anyone whose reason for choosing this is the agent angle is effectively choosing a Mac workflow.
- Credits meter AI, never exports
- Generating media and running models over footage consume credits; exporting does not, at any resolution the product supports, with no watermark and no quota. Cost therefore tracks how much analysis and generation a project needs rather than how much you edit or how often you publish. Monthly credits do not roll over, top-ups are available, and per-action cost depends on the model, its configuration and the duration of the media, so the same task is not a fixed price.
- Hosted AI routes inputs to third-party model providers
- The terms state that your inputs are licensed to the vendor to host, process and route to third-party AI providers for the purpose of running the service, and the privacy policy lists the sub-processors involved. Some of those providers retain outputs for safety review under their own terms. You retain rights in your inputs and own generated outputs to the extent copyright law allows, but the output is not guaranteed to be unique or free of third-party rights, and responsibility for the rights to what you upload or publish stays with you.
- It is an editing environment, not a distribution tool
- The vendor describes it as an integrated media environment for generating, editing and orchestrating media. Publishing to social platforms is not the product’s focus the way it is for scheduling-led editors, and there is no channel management, analytics or approval workflow. If your constraint is getting finished videos posted consistently across platforms, that is a different tool’s job even though the editing would work here.
- What it does not do
- It is not a collaborative review platform: there are no comments, approvals or shared brand templates, so agency workflows with sign-off are out. It does not run offline, since the hosted editor and AI features are cloud services. It does not put a visual editing surface over the composition layer, so code is the interface for precision work. There is no Windows desktop app yet, and no mobile app. And it does not decide what the edit should be — the agent can propose one, but the judgement remains editorial.