Models
Every model that can run shows its credit price up front. Unpriced models stay listed until their cost is verified — they just can't spend your credits yet.

Silero VAD
$0.00001/secondaudio-to-text
Detect speech presence and timestamps with accuracy and speed using the ultra-lightweight Silero VAD model

Cohere Transcribe
$0.00006944444/secondspeech-to-text
Cohere Transcribe turns your business audio into accurate text, ready for search, analytics, and automation

ACE Step
$0.0002/secondtext-to-audio
Generate music with lyrics from text using ACE-Step

ACE Step Audio Inpaint
$0.0002/secondaudio-to-audio
Modify a portion of provided audio with lyrics and/or style using ACE-Step

ACE Step Audio Outpaint
$0.0002/secondaudio-to-audio
Extend the beginning or end of provided audio with lyrics and/or style using ACE-Step

ACE Step Audio To Audio
$0.0002/secondaudio-to-audio
Generate music from a lyrics and example audio using ACE-Step

ACE Step Prompt To Audio
$0.0002/secondtext-to-audio
Generate music from a simple prompt using ACE-Step

FFmpeg API Compose
$0.0002/secondvideo-to-video
Compose videos from multiple media sources using FFmpeg API.

Ffmpeg Api
$0.0002/secondimage-to-image
ffmpeg endpoint for first, middle and last frame extraction from videos

Ffmpeg Api Merge Audio-Video
$0.0002/secondvideo-to-video
Merge videos with standalone audio files or audio from video files.

Flashvsr
$0.0005/megapixelvideo-to-video
Upscale your videos using FlashVSR with the fastest speeds!

DWPose Pose Prediction
$0.0006/secondimage-to-image
Predict poses from images.

DWPose Pose Prediction
$0.0006/secondvideo-to-video
Predict poses from videos.

Demucs
$0.0007/secondaudio-to-audio
SOTA stemming model for voice, drums, bass, guitar and more.

LTX-2 19B Distilled
$0.0008/megapixelimage-to-video
Generate video with audio from images using LTX-2 Distilled

LTX-2 19B Distilled
$0.0008/megapixeltext-to-video
Generate video with audio from text using LTX-2 Distilled

LTX-2 19B Distilled
$0.0008/megapixelaudio-to-video
Generate video with audio from audio, text and images using LTX-2 Distilled

LTX-2 19B Distilled
$0.0008/megapixelvideo-to-video
Extend videos with audio using LTX-2 Distilled

LTX-2 19B Distilled
$0.0008/megapixelvideo-to-video
Generate video with audio from videos using LTX-2 Distilled

Speech-To-text
$0.0008/secondspeech-to-text
Leverage the rapid processing capabilities of AI models to enable accurate and efficient real-time speech-to-text transcription.

Speech-to-Text
$0.0008/secondspeech-to-text
Leverage the rapid processing capabilities of AI models to enable accurate and efficient real-time speech-to-text transcription.

Speech-to-Text
$0.0008/secondspeech-to-text
Leverage the rapid processing capabilities of AI models to enable accurate and efficient real-time speech-to-text transcription.

Speech-to-Text
$0.0008/secondspeech-to-text
Leverage the rapid processing capabilities of AI models to enable accurate and efficient real-time speech-to-text transcription.

Video Upscaler
$0.0008/megapixelvideo-to-video
The video upscaler endpoint uses RealESRGAN on each frame of the input video to upscale the video to a higher resolution.

Ben-Video-Bg-Rm
$0.001/megapixelvideo-to-video
A model for high quality and smooth background removal for videos.

DDColor
$0.001/megapixelimage-to-image
Bring colors into old or new black and white photos with DDColor.

DeepFilterNet 3
$0.001/secondaudio-to-audio
Enhance speech audio by removing background noise and upsampling to 48KHz

LTX-2 19B Distilled
$0.001/megapixelimage-to-video
Generate video with audio from images using LTX-2 Distilled and custom LoRA

LTX-2 19B Distilled
$0.001/megapixeltext-to-video
Generate video with audio from text using LTX-2 Distilled and custom LoRA

LTX-2 19B Distilled
$0.001/megapixelvideo-to-video
Extend videos with audio using LTX-2 Distilled and custom LoRA

LTX-2 19B Distilled
$0.001/megapixelvideo-to-video
Generate video with audio from videos using LTX-2 Distilled and custom LoRA

LTX-2 19B Distilled
$0.001/megapixelaudio-to-video
Generate video with audio from audio, text and images using LTX-2 Distilled and custom LoRA

MMAudio V2
$0.001/secondvideo-to-video
MMAudio generates synchronized audio given video and/or text inputs. It can be combined with video models to get videos with audio.

MMAudio V2 Text to Audio
$0.001/secondtext-to-audio
MMAudio generates synchronized audio given text inputs. It can generate sounds described by a prompt.

NSFW Checker
$0.001/imagevision
Predict whether an image is NSFW or SFW.

NSFW Filter
$0.001/imagevision
Predict the probability of an image being NSFW.