Short-video creators or Chinese-literature teachers need to turn poem lines into imagery, calligraphy captions and music, then assemble a finished vertical video for account updates or classroom playback.
Manually assembling assets in editors like CapCut, or using separate image-generation, voice and caption tools and stitching the result by hand.
These videos have fixed aesthetic and format requirements; sourcing images, music and captions by hand is slow and stylistically inconsistent, while general generators do not understand poem shot structure or calligraphy layout.