Video clipping and repurposing system.
Overview
Three parts to the whole workflow. Doro Mind produces the video. Kobo builds the tool that turns it into clips. Doro Mind distributes. This proposal covers the middle.
Drop in a recording. Get back captioned, reframed clips with suggested copy, ranked and waiting at an approval gate. Doro Mind operates it, keeps shooting as they do today, and decides what posts, where, and when.
Sections 02 through 07 cover the build, which is the same under all three business models. Ownership, hosting, and whether Kobo stays involved are settled in section 08. Everything before it stays neutral.
Every figure carries a tag. Market is published 2026 pricing. Kobo est is built up from the scoped steps. Derived is arithmetic on figures Mimi Liu gave on the 7/30 call. Kobo's own numbers are still blank and render red.
The system
Seven stages: four bought, three built, all connected to run as a front to back pipeline.
The scoring layer and the approval gate are the parts nothing off the shelf does well enough here. Generic clip tools optimize for broad engagement, and none of them will accept a list of moments chosen elsewhere, so the choosing has to happen inside the system rather than be bought. A clip that performs well and misstates clozapine titration is a liability, not a win.
The pipeline does not care what kind of video it gets. A live panel, a recorded interview, a podcast appearance, all the same. So the video slate is not part of this proposal. Doro Mind decides what to make and the system handles what arrives.
Where the value is
None of these tools talk to each other. Scribe returns a transcript. Submagic wants a finished file. Metricool wants an asset and copy. Every handoff between them is glue that does not exist until someone writes it.
| Bought off the shelf | Built by Kobo |
|---|---|
| Speech to text | The wiring between the tools, including what happens when a step fails |
| Caption rendering and reframing | Scoring that ranks moments on whether they suit a caregiver, not on whether they might go viral |
| Scheduling and posting | Choosing which moments get cut, and cutting them |
| Analytics dashboards | The clinical term dictionary and risk flag list, built with Doro Mind clinicians |
| The approval queue, with rejection reasons captured rather than discarded | |
| Post copy drafted for each channel from what was actually said | |
| A record of what was rejected and what performed, which changes what gets cut next time |
If the deliverable were a list of subscriptions, Doro Mind would not need Kobo for it. The build is everything between the tools, plus two decisions no vendor will hand over: which moments are worth cutting, and which ones a clinician has to see first.
The durable assets
Three things hold their value as tools change. Who ends up holding them depends on the model.
| Asset | Why it keeps its value |
|---|---|
| The transcript layer | A searchable text record of every session. It is what makes clipping possible at all, and it stays useful whichever tools are in the stack. Held by Doro Mind under A and B, on Kobo systems under C. |
| The scoring criteria | A written definition of what a good Doro Mind clip is, agreed with clinicians. It is a document, so it is theirs under every model, and it works with any tool. |
| The sequence | Not tied to any one clipping vendor, so no tool provider can switch it off. Under A and B it runs on Doro Mind's servers and they own it outright. Under C it runs on Kobo's, and Kobo becomes the provider they depend on. |
Tool stack
Six pieces, each doing one job. Two bought, three built, one theirs.
Without something running the sequence, a person carries every file from one tool to the next, by hand, for every clip. That is the job the build removes.
ElevenLabs Scribe
BuyTurns each session into a timestamped, speaker separated transcript.
Fed the clinical dictionary, so clozapine, anosognosia, titration, and EASE come back spelled correctly. Generic models mangle all four. It reads the session video directly and returns timings word by word, which is what the scoring layer works from. Speaker separation matters because most sessions are panels.
Cutting
BuildTurns the ranked timecodes into clip files, vertical and square.
Built rather than bought, because every off the shelf clipper insists on choosing its own moments. Cutting from timecodes is a small piece of work and it is what keeps the scoring layer in charge.
Submagic
BuyBurns captions, strips filler and dead air, finishes per channel.
Captions are not optional here. A caregiver watching at 2am with the sound off still needs to follow it. It receives clips already cut, which keeps every file inside its size limits.
n8n
BuildRuns every step in order, on its own, so nobody carries files from one tool to the next.
Free to self host for internal use, so there are no per seat fees. What Kobo builds here is the sequence itself: what runs when, what happens when a step fails, and where each file goes next. Swapping a tool later means changing one step, not starting over.
LLM API
BuildScores clip candidates, flags clinical risk, drafts per channel copy.
The prompts and the scoring criteria are the custom work. The model itself can be swapped.
Metricool
TheirsSchedules approved clips and reports performance per clip.
Distribution is Doro Mind's. This is the suggested tool if they do not already have one, chosen because it has an API the pipeline can read performance back from.
Tools get replaced. If one raises its price or something better appears, one step changes and everything else keeps running. That is why the sequence sits in the middle, rather than betting on a single vendor to do the whole job.
The all-in-one question
Several products do transcription, clip selection, and captions in one subscription. Each was checked and set aside for the same reason: they decide which moments get cut, and none accepts a list of moments from somewhere else. Their ranking is tuned for broad engagement, which is the wrong question for a caregiver audience and no help on a dosing claim.
Phase 1 tests both against the same two archive sessions: moments chosen by the scoring criteria, and moments chosen by a vendor. If a vendor picks better, the cheaper path wins and the build shrinks. Worth running rather than assuming.
Whichever tools win, the built layers stay: the scoring, the clinical dictionary, the approval queue, and the record of what performed.
Pipeline, step by step
Thirteen steps. Ten built, two bought, one theirs.
Steps 05, 06, 10, and 13 are the ones nobody sells. A generic clipper does its own version of 03 through 08 and stops. It cannot tell whether a moment suits a caregiver in crisis, gives a clinician no way to catch a dosing claim before it goes out, and forgets everything between sessions. Step 07 is built for the same reason: every clipper insists on choosing the moments itself.
Division of labor
One lead per stage. Kobo builds, Doro Mind runs it day to day under all three models.
| Stage | Lead | Contributing |
|---|---|---|
| Pipeline architecture | Kobo | |
| Tool selection and trials | Kobo | Doro Mind approves the shortlist |
| Vendor terms for member footage | Kobo | Confirmed in writing per vendor. Doro Mind signs whatever is required. |
| Subscriptions and accounts | Doro Mind | Held by Doro Mind under A and B. Under C they sit with Kobo. |
| The sequence that runs the pipeline | Kobo | |
| Scoring layer | Kobo | Doro Mind defines what a good clip is |
| Clinical guardrails in scoring | Doro Mind | Kobo implements what they specify |
| Approval queue | Kobo | |
| Where the pipeline runs | Open | Their servers under A and B, Kobo systems under C. See section 08. |
| Connection to their scheduler | Kobo | The plumbing only. Doro Mind account access. |
| Distribution: what posts, where, and when | Doro Mind | Channel mix, timing, and final copy are theirs |
| Video production | Doro Mind | Unchanged. Patrick shoots and edits long form. |
| Loading sessions into the pipeline | Doro Mind | Every session, once the system is live |
| Approving clips | Doro Mind | Always. Nothing posts unreviewed. |
| Runbook and training | Kobo | |
| Named operator | Open | Needs a person and weekly hours under all three |
Phases
| Phase | What happens | Done when |
|---|---|---|
| Phase 1, Scope It | Tool trials against two real sessions from the archive. Scoring criteria defined with Doro Mind. Vendor terms for member footage confirmed in writing. Hosting decided. Costs confirmed. Phase 2 scope and price locked. | A working proof on real footage, and a fixed price for the build |
| Phase 2, Build It | The sequence, the scoring layer, the approval queue, the channel connections, and the existing archive run through from start to finish. | A session goes in, approved clips come out, published on schedule |
| Phase 3, Ship It | Runbook, operator training, readiness checklist. The Doro Mind operator runs one full cycle unaided. | Their operator runs the system without help. Under A and B, Kobo stops here. |
Phase 1 runs against real footage rather than a slide. If the tooling cannot handle their material, that surfaces in weeks and before any build budget is committed.
Three business models
The build above is the same in all three. What changes is when the money arrives, who ends up owning the system, and what Kobo has to keep running afterward.
Each panel below gives the flow, the money, and what the model means for each side. No recommendation is made here. Section 11 lists what each one has to clear.
Model A One-time Build and hand off $25K to $40K fixed Market
Fixed price per phase. Ownership transfers at handoff. Phase 1 alone runs $5K to $8K Market, and upkeep afterward is optional at $200 to $500 a month Market.
| For Doro Mind | For Kobo |
|---|---|
| Highest cash cost at the start, lowest cost over time. They own the system and hold the subscriptions, so member footage stays in their vendor relationships. | Simplest to sell and to close out. Clears every blocker in section 11. Revenue ends when the build does, and the code stops earning. |
Model B Financed Financed build $2K to $3K a month, 12 to 18 months Derived
Little or nothing down, then the same monthly fee for 12 to 18 months. The system runs on Doro Mind's servers from day one. Kobo retains admin until the final payment, then ownership and admin transfer.
| For Doro Mind | For Kobo |
|---|---|
| No capital outlay to start. They end up owning it, one year later than under A, and pay a financing premium for the delay. | A no-upfront close plus a year of predictable revenue, with none of C's licensing or custody work. Carries collection risk, so terms need a clause covering what happens to the system if payments stop. |
Model C Subscription Subscription on Kobo systems $2.5K to $3.5K a month, no end date Derived
Nothing upfront and a monthly fee with no end date. Doro Mind reaches the tool through Kobo and would be client one of a product. The blockers in section 11 all have fixes, and each is real work: owned code in place of n8n, data terms drafted, and a pitch built on low risk and cancel-anytime rather than ownership.
| For Doro Mind | For Kobo |
|---|---|
| Cheapest entry and no maintenance burden. They never own it, their footage sits on Kobo systems, and Kobo becomes a provider they depend on. Data terms and an exit path both have to exist before this can be offered. | The highest monthly rate of the three and the only one that keeps paying past the build. The engineering becomes a product asset that a second client can buy. In exchange Kobo takes on uptime, support, vendor breakage, and churn, and owes them a documented way to leave with their transcripts and criteria. |
C is the only model that changes what Kobo is. Priced at $3K a month, one client covers a meaningful share of a product line; three or four make it a business. The question is whether Doro Mind is client one of that product, which means the licensing and custody work happens now, or whether Kobo ships B, keeps the code, and generalizes it on its own schedule.
Money over time
The same build, three revenue shapes, and three obligation shapes underneath them.
| Model | Upfront | Monthly | Collected by month 24 | After month 24 | Kobo obligation |
|---|---|---|---|---|---|
| A. Build and hand off | $25K to $40K, split per phase | $0, or $200 to $500 upkeep | $25K to $52K | Upkeep only, if retained | Ends at handoff |
| B. Financed build | $0 or a small deposit | $2K to $3K for 12 to 18 months | $24K to $54K | Nothing. They own it. | Ends at transfer |
| C. Subscription | $0 | $2.5K to $3.5K, no end date | $60K to $84K | Continues, minus Kobo's hosting, license, and support costs | Uptime, support, vendor breakage, hosting, churn. No end date. |
All bands market-derived pending TBD: blended day rate. C collects the most and keeps collecting. The red band shows what that costs: A and B stop, C runs past the edge of the chart. Both lines have to be priced, and the monthly rate under C is what pays for the second one.
Numbers
Build estimate
31 to 50 working days. Kobo est
| Work block | Days |
|---|---|
| The sequence: wiring the tools, failure handling, file routing | 8 to 12 |
| Cutting: ranked timecodes to clip files | 2 to 4 |
| Scoring layer: prompts, criteria workshops, iteration on real sessions | 5 to 8 |
| Clinical dictionary and risk flags, with Doro Mind clinicians | 3 to 5 |
| Approval queue with rejection capture | 5 to 8 |
| Scheduler connection and performance readback | 2 to 4 |
| Archive backfill and end to end testing | 3 to 5 |
| Runbook, training, readiness checklist | 3 to 4 |
Elapsed: Phase 1 in 2 to 3 weeks, Phase 2 in 5 to 7 weeks. Inside the 12 week model with room. Model C adds 5 to 10 days to replace n8n with owned code, or an unpublished license fee to keep it.
Running costs sit well below the build. Transcription is priced by the audio hour, so a 90 minute session with the clinical dictionary attached costs well under a dollar. The monthly figures that matter are the finishing subscription and the model calls, both confirmed in Phase 1 against real footage.
Market evidence
| Comparable | Range | Source |
|---|---|---|
| n8n build with complex orchestration, AI, and custom logic | $15K to $50K+ | FlowEngine agency tiers |
| Production n8n automation stack with AI features | $5K to $35K+ | AY Automate agency guide |
| Post-build support plans | $200 to $500 a month | FlowEngine |
| Managed clipping and repurposing agencies | $1K to $10K+ a month | Clipping Culture services review |
| Off the shelf AI clipping tools | $15 to $100 a month | Clipping Culture services review |
The last row sets the anchor. In the client's head a $2,500 subscription sits next to $50 tools, so the gap has to be carried by the parts no tool sells: the scoring, the clinical guardrails, and the approval queue. The second to last row is the other anchor. Doing this as a service runs $1K to $10K a month elsewhere, forever, which is what any of these models replaces.
Blockers
What stands in the way of each model, row by row.
| Issue | A | B | C |
|---|---|---|---|
| n8n license. Free only for internal use; charging access to it requires a commercial agreement. | Clear | Clear | Blocked until n8n is replaced or licensed |
| Member footage custody. Session and app video from a psychiatric care provider. | Clear, stays on their servers | Clear, stays on their servers | Open, data terms and likely a BAA conversation |
| The ownership pitch. Doro Mind ends up owning the system outright. | Holds | Holds, one year later | Gone, replaced by low risk and cancel-anytime |
| Firm model. Kobo sells fixed-scope engagements, not open-ended retainers. | Clear | Clear, fixed term | Departs from it. Indefinite by design, and a deliberate change of business model. |
| Operations after launch. Uptime, support, vendor API breakage. | Theirs | Theirs at transfer | Kobo's, indefinitely |
Doro Mind's return
What the system's cost looks like against their own numbers.
| Figure | Value | Basis |
|---|---|---|
| Stated target, net new paid subscriptions | 30 | Mimi Liu, 7/30 call |
| Stated target, new ARR | $600K, described as ambitious | Mimi Liu, 7/30 call |
| Current ARR across 120 to 130 families | roughly $1.5M | Mimi Liu, 7/30 call |
| ARR per current family | roughly $11.5K to $12.5K | Derived |
| Families needed to cover $2.5K a month | between 2 and 3 a year | Derived |
Safe to claim: two to three net new families cover the system's cost under any of the three models, against a stated target of 30. Unsafe to claim: that the system delivers the 30. Contact volume depends on how often Doro Mind feeds it, and that accountability sits with them by design.
Decision
Two inputs stand between this and a priced proposal.
- Pick the model. A and B clear every blocker and end clean. C earns more, indefinitely, and turns the engineering into a product Kobo can sell again. That is a decision about what Kobo is, not just about this deal.
- Set the blended day rate. It converts the market bands into Kobo prices the same afternoon.
Outstanding
- Blended day rate
- Phase prices and durations
- Hosting decision
- Named operator and weekly hours
- Confirmed subscription costs
- Termination clause terms, if B
- Exit and data export terms, if C
Pick the model first. Every open number means something different under each one.