Pipeline
Step 1
分镜故事板
Text-to-imageGPT-Image2
- 生成方式
- 文生图
Break [Topic] into a 45-second educational explainer video with 9 scenes of about 5 seconds each. Requirements: 1. Each scene communicates only one core information point. 2. The sequence of scenes must build progressively. 3. The structure should suit a Vox-style / encyclopedia-collage explainer video. 4. Every scene must be visually expressible rather than relying only on subtitles. 5. Output: scene title, scene content, voice-over copy, and key visual elements. Example topic: “The Qin Unification of China” 9 scenes: 1. The Seven Warring States 2. Military power as the foundation, with rival states divided 3. The six states brought under one rule as Qin unifies the realm 4. Standardizing the writing system 5. Standardizing currency 6. Standardizing weights and measures 7. Standardizing cart gauges and transport routes 8. A unified imperial order 9. An influence lasting more than 2,000 years
Step 2
Reference-to-video
Requires: Storyboard
Image-to-videoSeedDance 2.0
- 生成方式
- Reference-to-video
Use the reference image as the sole visual foundation to generate a 5-second, 9:16, silent image-to-video clip. Animation requirements: Keep the paper cutout / scrapbook / stop-motion collage style. Move all elements—people, maps, text cards, coins, cart tracks, arrows, and so on—as independent cutout pieces. Include stepped-frame motion, slight paper jitter, segmented movement, and sticker-like bouncing. Possible small actions: Arrows advance, routes extend, coins rotate, seals stamp down, bamboo slips unfold, weights calibrate, horse carts move forward in stepped frames, and modules activate one by one. Important: Do not shake the entire image chaotically. Keep the main composition and important text stable. Give each scene one clear small event instead of making everything merely float. Finally: Edit the 9 clips together → add AI voice-over → subtitles → BGM. Keep the overall rhythm to one knowledge point every 4–5 seconds.