What it is in one line
The consumer creation product of multimodal model company ZhiXiang Future (HiDream.ai): it squeezes the whole "text idea → storyboard → assets → finished film" pipeline into a chat box and claims anyone can cut video without editing skills. The July 2026 vivago R1 release moved the pitch from "generate 10-second clips" to "unlimited-duration video generation and editing."
Who built it
Beijing ZhiXiang Future (HiDream.ai), founded March 2023, HQ in Hefei. Founder and CEO Mei Tao, former VP of JD.com, former computer-vision researcher at Microsoft Research Asia for over a decade, foreign academician of the Canadian Academy of Engineering. The company's self-developed UiT native full-modality architecture underpins the HiDream-O1 image models, which rank at the top of the open-source list on Artificial Analysis.
Read: the founder's background shapes the playbook. Mei Tao is a research-type executive — the technology narrative (native full-modality, world models) is far clearer than the channel narrative. That explains why the feature list is enormous while "who actually uses this heavily" is left for others to guess.
What it actually does
- Text-to-video → self-developed 2.0 engine turns text into 10-second HD video, claimed 4K, multi-shot, synchronized sound, with AI-generated music to dodge copyright
- Text-to-image / image-to-video / reference-to-video → three asset-generation entry points with multi-condition control (text, image, pose)
- Image toolset → image fusion, style imitation, product photo editing, interior redesign/staging, OOTD styling, group-photo stitching
- Marketing-specific features → Story Ad Director (three images + natural language → narrative marketing video), Amazon listing/A+ content designer, social cover designer, one-click social post generation and publishing
- Long-chain creation (R1) → the full "understand task → plot story → write storyboard → generate assets → generate long video" chain atomized and orchestrated, claimed to support unlimited duration
- Community and templates → 300+ templates; creators earn cash rewards (up to $600 per template) for submissions
What old behavior it replaces
To make a marketing video or e-commerce ad used to mean hiring an agency or production team: script, storyboard, actors and a studio, editing, color grading, voiceover. A 30-second spot cost anywhere from a few thousand to hundreds of thousands of yuan, on a weeks-long cycle.
vivago wants to compress that pipeline into a chat box — upload product/scene/hero images, describe the idea in natural language, and get a narrative video. For e-commerce sellers and brands it replaces the "outsource to a production team" move; for ordinary users it replaces the "I can't do it, so I won't" default.
Read: "anyone can use it" needs a discount. Generating a single short clip is indeed one tap for anyone, but continuously producing usable material and turning AI video into an ad that can actually run still requires someone who understands shot language. The tool lowers the bar; judgment does not come down with it.
Business model
Free base quota + pay-per-use credit packs + Pro membership (unlocks all duration options and sound effects). On the community side, template cash rewards incentivize UGC, forming a loop where users supply templates and templates attract users.
Read: consumer subscriptions and credits are the visible revenue; B-side (40,000 enterprise clients, orders from film-industry shareholders) is the bulk of the income statement. The investor list mixes Huace Film & TV and Shanghai Film Group — "shareholders who are also customers" is what separates this company from Kling or PixVerse.
Hard numbers
- traffic board: 4.60M MAU, +17.10% MoM (June 2026); on overseas growth and global growth boards
- July 2026: 1.5 billion RMB Series C; three rounds in three months totaling over 2.1 billion RMB; post-money valuation past $1B (company claim; valuation details undisclosed)
- Company disclosures: serving 100+ countries and regions, 50M+ registered users, 40,000+ enterprise clients
- Q1 2026 signed revenue over 400M RMB, exceeding the full year 2025
- vivago R1 claims an 85% usable-output success rate (vendor self-reported)
Four-way read
| Dimension |
Call |
| Founder-product fit |
High: Mei Tao is a vision researcher; the UiT architecture genuinely ranks near the top on Artificial Analysis image benchmarks, and the technology narrative is self-consistent |
| Product insight |
Medium: the feature list is over-stuffed (a dozen-plus tools); "from idea to feature film" is a good tagline but the primary user is not decided |
| Execution quality |
Upper-middle: self-developed models plus the R1 long-chain orchestration — engineering completeness beyond typical API-wrapper tools |
| Timing |
Good: it rides the video-model funding window (Kling's $3B and PixVerse's ~3B RMB Series C landed in the same weeks), but the window is also where competition is fiercest |
The call
A company whose funding narrative is stronger than its product narrative, and whose product sits on the line between "worth watching" and "unproven."
On hard conditions it ranks within the lane: a genuinely self-developed model, 2.1B RMB raised in three months, film-industry capital on the cap table, and 400M+ RMB in Q1 signed revenue. By those alone it is more solid than most AI-video startups.
The product side has real problems. First, "the AI director anyone can use" targets everyone and therefore no one; the actual retained users are most likely e-commerce sellers and short-video practitioners, not "everyone." Second, 4.6M MAU against a claimed 50M registered users and a $1B valuation means a deep registration-to-activity funnel. Third, R1's "unlimited duration" and the 85% usable-rate are vendor claims with no third-party verification.
Read: what's worth copying here is not the product but the capital playbook — rebranding "video generation" as a "world model" narrative and binding film-industry capital to order flow. That's the standard route for Chinese AI-video companies at this stage, and anyone holding comparable chips can mirror the narrative cadence.
What to watch next
① Whether vivago R1 shows third-party proof of paid conversion and long-form delivery — is "unlimited duration" a launch claim or a shippable product? Look for publicly documented commercial short-drama or promo-film cases
② Whether the Huace/Shanghai Film projects actually land and deliver — the first samples of the "shareholder-as-customer" model
③ Whether next month's traffic board MAU stays in double digits — at 4.6M the base is still expanding fast; if growth collapses, paid acquisition has stopped
What you can take from it
Product logic: for growth on a creation tool, copy the template-cash mechanism — users upload templates and earn cash per use (up to $600), turning your most platform-savvy users into content suppliers. Tool-class UGC loops almost always converge on this; the difference is only whether you pay in cash or traffic, and cash works better in cold start.
Pricing structure: the free quota + credits + membership three-layer structure is standard for video tools and worth referencing; its template-reward split is undisclosed and not directly transferable.
Verdict
Worth watching, with water in it. The funding, the team, and the self-developed models are real; 4.6M MAU is thin for a $1B-valuation company, and the self-reported 50M users and 85% usable-rate need third-party checking. Track it as a specimen of capital-and-narrative strategy for video-model companies rather than as a product specimen.