Skip to main content
ArtEmotion
Audio

ElevenLabs Scribe v2

Credit-based pricing.

ElevenLabs Scribe v2 preview
Output
audio

Inputs

  • Reference audio (required)
LLM-ready

API & LLM schema

Exact request contract for this model. Agents can fetch it from /api/v1/models?id=fal-ai/elevenlabs/speech-to-text/scribe-v2.

POST/api/v1/generate12 fields · 2 required
FieldTypeRequirementContract
model_idconstantRequiredArtEmotion model identifier.
extraobjectOptionalModel-specific settings may also be nested here.
max_creditsnumberOptionalReject before submission if the estimated list price exceeds this cap. · Range: 1–…
webhook_urlstringOptionalFormat: uri
webhook_secretstringOptionalOptional model input.
folder_idstringOptionalOptional model input.
language_codestringOptionalLanguage of the audio. Auto-detect works well for most clips. · Allowed: , eng, spa, fra, deu, por, ita, hin, zho, jpn, kor, ara, rus, ind, nld, tur, pol, swe, fil, msa, ron, ukr, ell, ces, dan, fin, bul, hrv, slk, tam, vie, tha, heb, hun, nor, cat · Default:
diarizebooleanOptionalAnnotate which speaker is talking at each point. · Default: false
tag_audio_eventsbooleanOptionalTag non-speech sounds like laughter, applause, or music. · Default: true
keytermsarray<string>OptionalVocabulary hints; each term is at most 5 words. Adds a surcharge to transcription, included by /api/v1/estimate. · Items: 0–1000
audio_urlstringRequiredFormat: uri
durationnumberOptionalEstimate-only fallback when no media is supplied. Generation measures source bytes before reserving credits; client duration cannot lower the bill. · Default: 60 · Range: 1–7200
Minimal request example
{
  "model_id": "fal-ai/elevenlabs/speech-to-text/scribe-v2",
  "audio_url": "https://example.com/audio"
}
Raw JSON Schema
{
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "$id": "https://www.artemotion.ai/api/v1/models?id=fal-ai%2Felevenlabs%2Fspeech-to-text%2Fscribe-v2",
  "title": "ElevenLabs Scribe v2 generation request",
  "description": "Request body accepted by POST /api/v1/generate for fal-ai/elevenlabs/speech-to-text/scribe-v2.",
  "type": "object",
  "properties": {
    "model_id": {
      "type": "string",
      "const": "fal-ai/elevenlabs/speech-to-text/scribe-v2",
      "description": "ArtEmotion model identifier."
    },
    "extra": {
      "type": "object",
      "additionalProperties": true,
      "description": "Model-specific settings may also be nested here."
    },
    "max_credits": {
      "type": "number",
      "minimum": 1,
      "description": "Reject before submission if the estimated list price exceeds this cap."
    },
    "webhook_url": {
      "type": "string",
      "format": "uri",
      "maxLength": 2048
    },
    "webhook_secret": {
      "type": "string",
      "maxLength": 512
    },
    "folder_id": {
      "type": "string"
    },
    "language_code": {
      "title": "Language",
      "description": "Language of the audio. Auto-detect works well for most clips.",
      "default": "",
      "type": "string",
      "enum": [
        "",
        "eng",
        "spa",
        "fra",
        "deu",
        "por",
        "ita",
        "hin",
        "zho",
        "jpn",
        "kor",
        "ara",
        "rus",
        "ind",
        "nld",
        "tur",
        "pol",
        "swe",
        "fil",
        "msa",
        "ron",
        "ukr",
        "ell",
        "ces",
        "dan",
        "fin",
        "bul",
        "hrv",
        "slk",
        "tam",
        "vie",
        "tha",
        "heb",
        "hun",
        "nor",
        "cat"
      ]
    },
    "diarize": {
      "title": "Diarize",
      "description": "Annotate which speaker is talking at each point.",
      "default": false,
      "type": "boolean"
    },
    "tag_audio_events": {
      "title": "Tag Audio Events",
      "description": "Tag non-speech sounds like laughter, applause, or music.",
      "default": true,
      "type": "boolean"
    },
    "keyterms": {
      "title": "Key Terms",
      "description": "Vocabulary hints; each term is at most 5 words. Adds a surcharge to transcription, included by /api/v1/estimate.",
      "type": "array",
      "maxItems": 1000,
      "items": {
        "type": "string",
        "minLength": 1,
        "maxLength": 50
      }
    },
    "audio_url": {
      "type": "string",
      "format": "uri"
    },
    "duration": {
      "type": "number",
      "minimum": 1,
      "maximum": 7200,
      "default": 60,
      "description": "Estimate-only fallback when no media is supplied. Generation measures source bytes before reserving credits; client duration cannot lower the bill."
    }
  },
  "required": [
    "model_id",
    "audio_url"
  ],
  "additionalProperties": false
}

FAQ

How much does ElevenLabs Scribe v2 cost on ArtEmotion?

Credit-based pricing. You pay in ArtEmotion credits — every plan and top-up converts USD to credits at a fixed rate.

Do I get my credits back if ElevenLabs Scribe v2 fails?

Yes — failed generations are never charged. The credits are released back to your balance automatically.

Can I call ElevenLabs Scribe v2 from the API?

Yes. Use POST /api/v1/generate with model_id: "fal-ai/elevenlabs/speech-to-text/scribe-v2". See the API reference for the full schema.

Where are my generations stored?

Every output is saved to your personal Library. You can export or delete everything any time from Privacy & deletion.

Ready to generate with ElevenLabs Scribe v2?

Start now →

No commitment. New accounts get 50 free credits on signup. See pricing.