Models

Every model that can run shows its credit price up front. Unpriced models stay listed until their cost is verified — they just can't spend your credits yet.

  • Silero VAD

    $0.00001/second

    audio-to-text

    Detect speech presence and timestamps with accuracy and speed using the ultra-lightweight Silero VAD model

  • Cohere Transcribe

    $0.00006944444/second

    speech-to-text

    Cohere Transcribe turns your business audio into accurate text, ready for search, analytics, and automation

  • ACE Step

    $0.0002/second

    text-to-audio

    Generate music with lyrics from text using ACE-Step

  • ACE Step Audio Inpaint

    $0.0002/second

    audio-to-audio

    Modify a portion of provided audio with lyrics and/or style using ACE-Step

  • ACE Step Audio Outpaint

    $0.0002/second

    audio-to-audio

    Extend the beginning or end of provided audio with lyrics and/or style using ACE-Step

  • ACE Step Audio To Audio

    $0.0002/second

    audio-to-audio

    Generate music from a lyrics and example audio using ACE-Step

  • ACE Step Prompt To Audio

    $0.0002/second

    text-to-audio

    Generate music from a simple prompt using ACE-Step

  • FFmpeg API Compose

    $0.0002/second

    video-to-video

    Compose videos from multiple media sources using FFmpeg API.

  • Ffmpeg Api

    $0.0002/second

    image-to-image

    ffmpeg endpoint for first, middle and last frame extraction from videos

  • Ffmpeg Api Merge Audio-Video

    $0.0002/second

    video-to-video

    Merge videos with standalone audio files or audio from video files.

  • Flashvsr

    $0.0005/megapixel

    video-to-video

    Upscale your videos using FlashVSR with the fastest speeds!

  • DWPose Pose Prediction

    $0.0006/second

    image-to-image

    Predict poses from images.

  • DWPose Pose Prediction

    $0.0006/second

    video-to-video

    Predict poses from videos.

  • Demucs

    $0.0007/second

    audio-to-audio

    SOTA stemming model for voice, drums, bass, guitar and more.

  • LTX-2 19B Distilled

    $0.0008/megapixel

    image-to-video

    Generate video with audio from images using LTX-2 Distilled

  • LTX-2 19B Distilled

    $0.0008/megapixel

    text-to-video

    Generate video with audio from text using LTX-2 Distilled

  • LTX-2 19B Distilled

    $0.0008/megapixel

    audio-to-video

    Generate video with audio from audio, text and images using LTX-2 Distilled

  • LTX-2 19B Distilled

    $0.0008/megapixel

    video-to-video

    Extend videos with audio using LTX-2 Distilled

  • LTX-2 19B Distilled

    $0.0008/megapixel

    video-to-video

    Generate video with audio from videos using LTX-2 Distilled

  • Speech-To-text

    $0.0008/second

    speech-to-text

    Leverage the rapid processing capabilities of AI models to enable accurate and efficient real-time speech-to-text transcription.

  • Speech-to-Text

    $0.0008/second

    speech-to-text

    Leverage the rapid processing capabilities of AI models to enable accurate and efficient real-time speech-to-text transcription.

  • Speech-to-Text

    $0.0008/second

    speech-to-text

    Leverage the rapid processing capabilities of AI models to enable accurate and efficient real-time speech-to-text transcription.

  • Speech-to-Text

    $0.0008/second

    speech-to-text

    Leverage the rapid processing capabilities of AI models to enable accurate and efficient real-time speech-to-text transcription.

  • Video Upscaler

    $0.0008/megapixel

    video-to-video

    The video upscaler endpoint uses RealESRGAN on each frame of the input video to upscale the video to a higher resolution.

  • Ben-Video-Bg-Rm

    $0.001/megapixel

    video-to-video

    A model for high quality and smooth background removal for videos.

  • DDColor

    $0.001/megapixel

    image-to-image

    Bring colors into old or new black and white photos with DDColor.

  • DeepFilterNet 3

    $0.001/second

    audio-to-audio

    Enhance speech audio by removing background noise and upsampling to 48KHz

  • LTX-2 19B Distilled

    $0.001/megapixel

    image-to-video

    Generate video with audio from images using LTX-2 Distilled and custom LoRA

  • LTX-2 19B Distilled

    $0.001/megapixel

    text-to-video

    Generate video with audio from text using LTX-2 Distilled and custom LoRA

  • LTX-2 19B Distilled

    $0.001/megapixel

    video-to-video

    Extend videos with audio using LTX-2 Distilled and custom LoRA

  • LTX-2 19B Distilled

    $0.001/megapixel

    video-to-video

    Generate video with audio from videos using LTX-2 Distilled and custom LoRA

  • LTX-2 19B Distilled

    $0.001/megapixel

    audio-to-video

    Generate video with audio from audio, text and images using LTX-2 Distilled and custom LoRA

  • MMAudio V2

    $0.001/second

    video-to-video

    MMAudio generates synchronized audio given video and/or text inputs. It can be combined with video models to get videos with audio.

  • MMAudio V2 Text to Audio

    $0.001/second

    text-to-audio

    MMAudio generates synchronized audio given text inputs. It can generate sounds described by a prompt.

  • NSFW Checker

    $0.001/image

    vision

    Predict whether an image is NSFW or SFW.

  • NSFW Filter

    $0.001/image

    vision

    Predict the probability of an image being NSFW.