
Output
audio
Example prompt
A slow cinematic orchestral swell — strings building from pianissimo to fortissimo, low brass entering at the climax, kettle drums rolling in the backgroundTry this prompt →
LLM-ready
API & LLM schema
Exact request contract for this model. Agents can fetch it from /api/v1/models?id=fal-ai/stable-audio-3/medium/text-to-audio.
POST
/api/v1/generate16 fields · 2 required| Field | Type | Requirement | Contract |
|---|---|---|---|
model_id | constant | Required | ArtEmotion model identifier. |
extra | object | Optional | Model-specific settings may also be nested here. |
max_credits | number | Optional | Reject before submission if the estimated list price exceeds this cap. · Range: 1–… |
webhook_url | string | Optional | Format: uri |
webhook_secret | string | Optional | Optional model input. |
folder_id | string | Optional | Optional model input. |
prompt | string | Required | Optional model input. |
duration | number | Optional | Duration of the generated audio in seconds (1–380). · Default: 30 · Range: 1–380 |
output_format | string | Optional | Audio container/codec for the output file. · Allowed: mp3, wav, flac, ogg, opus, m4a, aac · Default: mp3 |
bitrate | string | Optional | Bitrate for compressed output formats. · Allowed: 192k, 320k · Default: 192k |
negative_prompt | string | Optional | Describe what the audio should avoid. · Default: |
num_inference_steps | number | Optional | Number of diffusion steps. Higher = slightly better quality, slower. · Default: 8 · Range: 1–100 |
guidance_scale | number | Optional | How closely the model follows the prompt. Higher = more literal. · Default: 1 · Range: 0–25 |
seed | integer | Optional | Set a seed for reproducible results. Leave blank for random. |
enable_prompt_expansion | boolean | Optional | Let the model expand short prompts into richer descriptions. · Default: false |
enable_safety_checker | boolean | Optional | Screen the output for unsafe content. · Default: true |
Minimal request example
{
"model_id": "fal-ai/stable-audio-3/medium/text-to-audio",
"prompt": "A slow cinematic orchestral swell — strings building from pianissimo to fortissimo, low brass entering at the climax, kettle drums rolling in the background"
}Raw JSON Schema
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"$id": "https://www.artemotion.ai/api/v1/models?id=fal-ai%2Fstable-audio-3%2Fmedium%2Ftext-to-audio",
"title": "Stable Audio 3 generation request",
"description": "Request body accepted by POST /api/v1/generate for fal-ai/stable-audio-3/medium/text-to-audio.",
"type": "object",
"properties": {
"model_id": {
"type": "string",
"const": "fal-ai/stable-audio-3/medium/text-to-audio",
"description": "ArtEmotion model identifier."
},
"extra": {
"type": "object",
"additionalProperties": true,
"description": "Model-specific settings may also be nested here."
},
"max_credits": {
"type": "number",
"minimum": 1,
"description": "Reject before submission if the estimated list price exceeds this cap."
},
"webhook_url": {
"type": "string",
"format": "uri",
"maxLength": 2048
},
"webhook_secret": {
"type": "string",
"maxLength": 512
},
"folder_id": {
"type": "string"
},
"prompt": {
"type": "string"
},
"duration": {
"title": "Duration (s)",
"description": "Duration of the generated audio in seconds (1–380).",
"default": 30,
"type": "number",
"minimum": 1,
"maximum": 380,
"multipleOf": 1
},
"output_format": {
"title": "Output Format",
"description": "Audio container/codec for the output file.",
"default": "mp3",
"type": "string",
"enum": [
"mp3",
"wav",
"flac",
"ogg",
"opus",
"m4a",
"aac"
]
},
"bitrate": {
"title": "Bitrate",
"description": "Bitrate for compressed output formats.",
"default": "192k",
"type": "string",
"enum": [
"192k",
"320k"
]
},
"negative_prompt": {
"title": "Negative Prompt",
"description": "Describe what the audio should avoid.",
"default": "",
"type": "string"
},
"num_inference_steps": {
"title": "Inference Steps",
"description": "Number of diffusion steps. Higher = slightly better quality, slower.",
"default": 8,
"type": "number",
"minimum": 1,
"maximum": 100,
"multipleOf": 1
},
"guidance_scale": {
"title": "Guidance Scale",
"description": "How closely the model follows the prompt. Higher = more literal.",
"default": 1,
"type": "number",
"minimum": 0,
"maximum": 25,
"multipleOf": 0.5
},
"seed": {
"title": "Seed",
"description": "Set a seed for reproducible results. Leave blank for random.",
"type": "integer"
},
"enable_prompt_expansion": {
"title": "Prompt Expansion",
"description": "Let the model expand short prompts into richer descriptions.",
"default": false,
"type": "boolean"
},
"enable_safety_checker": {
"title": "Safety Checker",
"description": "Screen the output for unsafe content.",
"default": true,
"type": "boolean"
}
},
"required": [
"model_id",
"prompt"
],
"additionalProperties": false
}FAQ
How much does Stable Audio 3 cost on ArtEmotion?
Credit-based pricing. You pay in ArtEmotion credits — every plan and top-up converts USD to credits at a fixed rate.
Do I get my credits back if Stable Audio 3 fails?
Yes — failed generations are never charged. The credits are released back to your balance automatically.
Can I call Stable Audio 3 from the API?
Yes. Use POST /api/v1/generate with model_id: "fal-ai/stable-audio-3/medium/text-to-audio". See the API reference for the full schema.
Where are my generations stored?
Every output is saved to your personal Library. You can export or delete everything any time from Privacy & deletion.
Ready to generate with Stable Audio 3?
Start now →No commitment. New accounts get 50 free credits on signup. See pricing.