NSFW Text to Video: Generating Without a Reference Photo
An nsfw text to video run has no photo to reject, which changes where the refusals come from. Here is what the prompt has to survive and what a second costs.
From $0.07 per second, 480p generation only

Why nsfw text to video behaves differently
With nsfw text to video there is no upload to block, so the moderation that stops image workflows never fires. The refusals move to the prompt and to the rendered output instead.
Image workflows die at upload when a person is detected. An nsfw text to video request carries no reference, so that entire class of rejection does not apply.
Every nsfw text to video job requires a prompt, up to 20,000 characters. Reference-driven runs can omit it; a pure text run cannot, and neither can any request with content_filter set false.
The prompt itself, and the rendered output. An nsfw text to video job that passes submission can still fail minutes later when the finished frames are reviewed.
Without a reference image, an nsfw text to video result will not hold the same character across separate generations. Describing a person in words gets you a type, not a specific face.
Speech, effects and music render with the picture. Quoted lines in an nsfw text to video prompt come back as generated dialogue rather than silence.
No generator here. Nsfw text to video runs happen on the hosted API, billed per second of output, and every figure below comes from published parameters.
What nsfw text to video gives you that image workflows do not
Four differences worth knowing before choosing a mode.
Nothing to upload
No hosting, no public URLs, no format checks. An nsfw text to video request is a single JSON body, which removes the most common cause of submission errors.
Full aspect ratio choice
All six ratios plus adaptive. Frame-driven runs are locked to adaptive, so nsfw text to video is the only mode where you pick the shape freely.
Thirty seconds, multiple shots
One nsfw text to video generation runs to 30 seconds and can contain several shots. Rivals stop at 15.
Cheapest billing path
No reference video means no extra billed seconds. An nsfw text to video clip is charged on output length alone, which makes it the least expensive mode per finished second.
Four steps for nsfw text to video
- 1
Write the scene
Describe subject, motion and camera in one prompt. An nsfw text to video model reads the whole thing, so ordering matters less than specificity.
- 2
Add dialogue if you want it
Put spoken lines in double quotes. The nsfw text to video pass generates that speech in sync rather than leaving a silent track.
- 3
Pick shape, length, resolution
Any of six ratios, four to thirty seconds, at 480p, 720p or 1080p. No nsfw text to video request reaches 4K on this model.
- 4
Submit and poll
Two to five minutes. Polling is free and a failed nsfw text to video job is refunded in full.
Nsfw text to video, answered
Differently restricted. There is no reference photo to reject, so the face-detection failures that end most image runs never happen. In exchange the prompt carries all the risk, and a job can still fail after generation when the output is reviewed.
Always. Text is the only input, so the prompt is required and can run to 20,000 characters. It is also required on any request that sets content_filter to false, regardless of mode, which catches people out when they switch channels.
Per second of output, from $0.118589 at 480p to $0.461856 at 1080p. Without a reference video there is no extra billed footage, so a ten-second 1080p clip is about $4.62 and a five-second 480p clip about $0.59. Failed jobs are refunded.
Not across separate generations. Words describe a type rather than a face, so each run produces a different person. If consistency matters, supply reference images instead and accept that the upload will be reviewed.
Error 80006 is a safety refusal, and it deliberately does not say which stage fired. A synchronous 400 is different: that is schema validation, usually a bad resolution value or a missing prompt, and nothing was charged. Check whether a task id came back before assuming moderation.
It routes the request to a less restrictive channel at the same price, and it makes the prompt mandatory. The official documentation calls it a routing control rather than an off switch, and upstream policy checks still run on every nsfw text to video job.
Named real celebrities, third-party intellectual property and illegal content. Describing a celebrity in words rather than uploading a photo does not get around it, because the block sits inside the model. A refused job is refunded rather than partially delivered.
Twenty thousand characters, which is far more than most scenes need. Length past a few hundred words tends to dilute rather than sharpen the result, so specificity beats volume.
Yes, but it stops being a text-only run. Adding video_urls turns it into a reference job, which brings back the review on uploaded material and adds the source clip's length to the bill. Any request containing video also carries a floor of five thirds of the output duration, so a two-second source with a five-second output bills nine seconds.
Run nsfw text to video without an upload
One prompt, no reference hosting, six aspect ratios, per-second rates and refunds when an nsfw text to video job fails.
Related: Seedance 2.5 API