MODEL CATALOG

Models for every workload

Compare capabilities, pricing, and providers across one unified catalog.
Available
1195
Providers
92
Vendors92
1195

Explore models

Page 1 of 24

LLM

deepseek/deepseek-v4-flash

DeepSeek

DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fa...

In: $0.14Out: $0.28/1M tokens
1M context
IMAGE

google/nano-banana-2/edit

Google

Nano Banana 2 is Google's new state-of-the-art image generation and editing model

$0.06 / call
VIDEO

xai/grok-imagine-video/1.5/image-to-video

xAI

Grok Imagine Video 1.5 is xAI's image-to-video AI model. Bring static images to life with fluid animation, consistent character motion, and professional-grade video output.

$0.40 / call
LLM

xiaomi/mimo-v2.5

Xiaomi

MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in multimodal perception across image and v...

In: $0.40Out: $2.00/1M tokens
1M context
IMAGE

openai/gpt-image-2/edit

OpenAI

GPT Image 2 Editing supports image editing and multi-image synthesis with high-quality results.

$0.20 / call
LLM

minimax/minimax-m3

Minimax

MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding, an...

In: $0.30Out: $1.20/1M tokens
1M context
IMAGE

google/nano-banana-pro/edit

Google

Nano Banana Pro is Google's new state-of-the-art image generation and editing model

$0.12 / call
LLM

tencent/hy3-preview

Tencent

Hy3 preview is a high-efficiency Mixture-of-Experts model from Tencent designed for agentic workflows and production use. It supports configurable reasoning levels across disabled, low, and high modes...

In: $0.07Out: $0.26/1M tokens
262.1K context
IMAGE

google/nano-banana-2

Google

Nano Banana 2 is Google's new state-of-the-art fast image generation and editing model

$0.06 / call
LLM

anthropic/claude-opus-4.7

Anthropic

Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on c...

In: $5.00Out: $25.00/1M tokens
1M context
IMAGE

openai/gpt-image-2

OpenAI

GPT Image 2, OpenAI's latest image model, is capable of making fine-grained, detailed edits to images.

$0.20 / call
LLM

deepseek/deepseek-v4-pro

DeepSeek

DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reaso...

In: $0.43Out: $0.87/1M tokens
1M context
LLM

anthropic/claude-opus-4.8

Anthropic

Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token context windo...

In: $5.00Out: $25.00/1M tokens
1M context
IMAGE

google/nano-banana-pro

Google

Nano Banana Pro is Google's new state-of-the-art image generation and editing model

$0.12 / call
LLM

z-ai/glm-5.2

Z-Ai

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering...

In: $1.40Out: $4.40/1M tokens
1M context
VIDEO

kwaivgi/kling-video/v3/pro/image-to-video

KwaiVGI

Kling Video V3 Pro is KwaiVGI's image-to-video AI model. Bring static images to life with fluid animation, consistent character motion, and professional-grade video output.

$0.84 / call
IMAGE

google/nano-banana/edit

Google

Nano Banana Pro is Google's new state-of-the-art image generation and editing model

$0.15 / call
LLM

anthropic/claude-sonnet-4.6

Anthropic

Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative development, complex codebase navigation, ...

In: $3.00Out: $15.00/1M tokens
1M context
LLM

anthropic/claude-sonnet-5

Anthropic

Claude Sonnet 5 is a Sonnet-class model for high-quality coding, agentic workflows, reasoning, vision, structured outputs, and tool use.

In: $2.00Out: $10.00/1M tokens
1M context
VIDEO

bytedance/seedance/2.0/image-to-video

Bytedance

ByteDance's most advanced image-to-video model transforming still images into cinematic video with native audio, multi-shot editing, and director-level camera control for professional-grade video cre...

$1.00 / call
LLM

stepfun/step-3.7-flash

Stepfun

Step 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model. It pairs a 196B-parameter language backbone with a vision encoder for native image and video understanding, acti...

In: $0.20Out: $1.15/1M tokens
256K context
IMAGE

bytedance/seedream/4.5/edit

Bytedance

A new-generation image creation model from ByteDance, Seedream 4.5 integrates text-to-image generation and image editing into a single unified architecture, delivering high-fidelity visuals, precise p...

$0.04 / call
LLM

openai/gpt-5.5

OpenAI

GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved token efficiency on hard tasks. It feature...

In: $5.00Out: $30.00/1M tokens
1.1M context
LLM

google/gemini-3-flash-preview

Google

Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers near Pro level reasoning and tool use performance ...

In: $0.50Out: $3.00/1M tokens
1M context
IMAGE

bfl/flux-2/pro

BFL

Flux 2 Pro by BFL - generate stunning images from text prompts with state-of-the-art AI. Supports multiple aspect ratios, styles, and high-resolution output for creative and commercial use.

$0.03 / call
LLM

deepseek/deepseek-v3.2

DeepSeek

DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use performance. It introduces DeepSeek Sparse Attention (DSA), a fin...

In: $0.25Out: $0.38/1M tokens
131.1K context
IMAGE

bfl/flux-1.1/pro

BFL

Flux 1.1 Pro by BFL - generate stunning images from text prompts with state-of-the-art AI. Supports multiple aspect ratios, styles, and high-resolution output for creative and commercial use.

$0.04 / call
LLM

z-ai/glm-5.1

Z-Ai

GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around minute-level interactions, GLM-5.1 can work ...

In: $1.05Out: $3.50/1M tokens
202.8K context
VIDEO

bytedance/seedance/2.0/reference-to-video

Bytedance

ByteDance's most advanced reference-to-video model generating cinematic video guided by reference content, with native audio, multi-shot editing, and director-level camera control for professional-gr...

$0.60 / call
LLM

google/gemini-2.5-flash-lite

Google

Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It offers improved throughput, faster token generation, and better...

In: $0.10Out: $0.40/1M tokens
1M context
IMAGE

bfl/flux-1/kontext

BFL

Flux 1 Kontext by BFL - AI-powered image editing, style transfer, and transformation. Edit photos with natural language instructions, remove backgrounds, change styles, and enhance images effortlessly...

$0.04 / call
IMAGE

mirelo/sfx1.6/video-to-video

Mirelo

Sfx1.6 Video To Video is Mirelo's video-to-video AI model. Transform, enhance, and edit video content using text prompts - from style changes to object manipulation and scene modification.

$0.0013 / call
LLM

google/gemini-2.5-flash

Google

Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It includes built-in "thinking" capabilities, en...

In: $0.30Out: $2.50/1M tokens
1M context
IMAGE

google/nano-banana

Google

Google's famous original image generation and editing model.

$0.04 / call
IMAGE

bfl/flux-2/pro/edit

BFL

Flux 2 Pro Edit is BFL's intelligent image editing model. Transform, retouch, and reimagine existing images using text prompts - from background replacement to artistic style conversion.

$0.03 / call
AUDIO

mirelo/sfx1.6/text-to-audio

Mirelo

Sfx1.6 by Mirelo - generate music, sound effects, and audio from text descriptions with AI. Create original compositions, ambient sounds, and audio content for any creative project.

Pricing on request
LLM

moonshotai/kimi-k2.6

Moonshotai

Kimi K2.6 is Moonshot AI's next-generation multimodal model, designed for long-horizon coding, coding-driven UI/UX generation, and multi-agent orchestration. It handles complex end-to-end coding tasks...

In: $0.75Out: $3.50/1M tokens
262.1K context
LLM

moonshotai/kimi-k3

Moonshotai

Kimi K3 is Moonshot AI's latest generation model with advanced coding, reasoning, and multi-agent capabilities.

In: $3.00Out: $15.00/1M tokens
1M context
IMAGE

mirelo/sfx1.6/inpaint-audio

Mirelo

Sfx1.6 Inpaint Audio by Mirelo - advanced AI model for audio-to-audio. Delivers high-quality results with fast inference, suitable for both creative and production workflows.

$0.0013 / call
IMAGE

mirelo/sfx1.6/extend-audio

Mirelo

Sfx1.6 Extend Audio by Mirelo - advanced AI model for audio-to-audio. Delivers high-quality results with fast inference, suitable for both creative and production workflows.

Pricing on request
IMAGE

xai/grok-imagine-image/edit

xAI

Grok Imagine Image Edit is xAI's intelligent image editing model. Transform, retouch, and reimagine existing images using text prompts - from background replacement to artistic style conversion.

$0.02 / call
VIDEO

kwaivgi/kling-video/v2.5-turbo/pro/image-to-video

KwaiVGI

Kling Video V2.5 Turbo Pro is KwaiVGI's image-to-video AI model. Bring static images to life with fluid animation, consistent character motion, and professional-grade video output.

$0.35 / call
LLM

xiaomi/mimo-v2.5-pro

Xiaomi

MiMo-V2.5-Pro is Xiaomi’s flagship model, delivering strong performance in general agentic capabilities, complex software engineering, and long-horizon tasks, with top rankings on benchmarks such as C...

In: $1.00Out: $3.00/1M tokens
1M context
LLM

openai/gpt-4o-mini

OpenAI

GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more ...

In: $0.15Out: $0.60/1M tokens
128K context
IMAGE

bytedance/seedvr/upscale/image

Bytedance

Seedvr Upscale Image by Bytedance - AI-powered image editing, style transfer, and transformation. Edit photos with natural language instructions, remove backgrounds, change styles, and enhance images ...

$0.0010 / call
LLM

google/gemini-3.1-flash-lite

Google

Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, and PDF inputs, and is designed for light...

In: $0.25Out: $1.50/1M tokens
1M context
IMAGE

bfl/flux-1.1/pro-ultra

BFL

Flux 1.1 Pro Ultra by BFL - generate stunning images from text prompts with state-of-the-art AI. Supports multiple aspect ratios, styles, and high-resolution output for creative and commercial use.

$0.06 / call
LLM

google/gemini-3.5-flash

Google

Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel age...

In: $1.50Out: $9.00/1M tokens
1M context
IMAGE

bfl/flux-1/lora

BFL

Flux 1 Lora is BFL's advanced text-to-image AI model. Create photorealistic images, illustrations, and concept art from natural language descriptions with exceptional detail and prompt adherence.

$0.04 / call
LLM

openai/gpt-5.4

OpenAI

GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window (922K input, 128K output) with support for text and image input...

In: $2.50Out: $15.00/1M tokens
1.1M context