| Google | | TextText generationImage understanding | Fast drafting, summaries, and routine multimodal tasks. | - Responsive text generation with image-reference support.
| - Less suited than Pro models to long, highly complex reasoning.
|
| Google | | TextText generationImage understanding | Detailed writing, analysis, and complex instructions. | - Stronger reasoning and instruction handling with image context.
| - Typically slower and more credit-intensive than Flash.
|
| Google | | TextText generationImage understanding | Rapid ideation and high-volume text workflows. | - Balances speed with current-generation multimodal understanding.
| - Prioritizes response speed over maximum analytical depth.
|
| Google | | TextText generationImage understanding | Complex creative briefs and structured production analysis. | - High-capability reasoning across text and visual references.
| - Heavier jobs can take longer than Flash variants.
|
| Google | | TextText generationImage understanding | The most demanding planning, reasoning, and writing tasks. | - Our most capable Gemini option for nuanced instructions.
| - Best reserved for work where added quality justifies extra time and credits.
|
| Anthropic | | TextText generationImage understanding | Demanding reasoning, agentic coding, and professional knowledge work. | - Handles complex, long-running tasks and supports tool calling.
| - Uses more credits than lighter text models.
|
| Anthropic | | TextText generationImage understanding | Long-running coding, vision reasoning, and multi-step knowledge work. | - Strong planning and verification across extended tasks.
| - The connected model does not currently support tool calling or Claude Code.
|
| OpenAI | | TextText generationImage understanding | Fast, high-volume text and multimodal workflows. | - Lowest-cost GPT-5.6 tier with configurable reasoning effort.
| - Prioritizes speed and efficiency over maximum capability.
|
| OpenAI | | TextText generationImage understanding | Balanced everyday coding and professional work. | - Combines stronger reasoning with practical cost and latency.
| - Less capable than Sol on the most demanding tasks.
|
| OpenAI | | TextText generationImage understanding | Frontier reasoning, complex coding, and high-autonomy tasks. | - The highest-capability GPT-5.6 tier for difficult work.
| - Uses more credits and can take longer than Luna or Terra.
|
| Google | | ImageText to imageImage to image | Reference-led image generation and iterative visual editing. | - Supports generation or editing with up to 14 reference images.
| - Dense multi-reference prompts can still require iteration.
|
| OpenAI | | ImageText to imageImage to image | Instruction-accurate image creation, text rendering, and edits. | - Flexible generation and editing with broad reference support.
| - 4K output uses more credits than 1K or 2K.
|
| xAI | | | Quick image concepts from straightforward prompts. | - Simple text-to-image workflow with useful aspect-ratio choices.
| - Fewer detailed controls than reference-led image models.
|
| xAI | | | Quick, prompt-guided changes to a single image. | - A focused image-editing workflow with simple setup.
| - Accepts one source image and offers fewer output controls than broader editing models.
|
| Alibaba | | ImageText to imageImage to image | Flexible image generation and editing with multiple references. | - Supports up to nine references and wide aspect ratios.
| - Reference-based jobs are capped below the 4K text-only mode.
|
| Alibaba | | ImageText to imageImage to image | Higher-quality Wan image generation and editing. | - Pro quality tier with multi-reference and resolution controls.
| - Reference-based jobs are capped below the 4K text-only mode.
|
| Alibaba | | ImageText to imageImage to image | Efficient image generation and editing with multilingual typography. | - Supports 1K/2K output and edits with up to 3 source images.
| - Initial editing support accepts JPG/JPEG, PNG, WebP, BMP, and GIF source files; TIFF is not supported.
|
| Alibaba | | ImageText to imageImage to image | Higher-fidelity layouts, typography, and reference-led image editing. | - Pro-quality 1K/2K output with up to 3 source images for editing.
| - Initial editing support accepts JPG/JPEG, PNG, WebP, BMP, and GIF source files; TIFF is not supported.
|
| ByteDance | | ImageText to imageImage to image | Everyday image generation and reference-led edits. | - Offers High and Basic quality modes with up to 14 reference images.
| - TUBAGEN does not expose separate resolution or output-format controls for this option.
|
| ByteDance | | ImageText to imageImage to image | Polished image generation and controlled image-to-image work. | - High-quality generation with up to ten input images.
| - Fixed PNG output and a heavier credit cost than Lite.
|
| Google | | ImageText to imageImage to image | Consistent image edits guided by several references. | - Supports up to eight reference images and common aspect ratios.
| - Fewer reference slots than Nano Banana 2.
|
| Google | | | Simple text-to-image generation in varied aspect ratios. | - Straightforward, efficient prompt-to-image workflow.
| - Does not accept reference-image input in TUBAGEN.
|
| Google | | | Prompt-guided edits that combine several source images. | - Accepts up to ten source images and supports common aspect ratios.
| - Requires an input image and is intended for editing rather than text-only generation.
|
| Alibaba | | ImageText to imageImage to image | Fast concepts using a small reference set. | - Supports up to three references and a broad set of aspect ratios.
| - Less suitable for jobs needing large reference collections.
|
| Black Forest Labs | | ImageText to imageImage to image | Detailed commercial visuals and images containing typography. | - Strong detail with generation or editing from up to eight references.
| - Complex layouts and exact text may still need retries.
|
| Topaz Labs | | | Upscaling a single image to twice its source resolution. | - A dedicated image-enhancement tool that needs no prompt.
| - Requires one source image and currently uses a fixed 2× scale.
|
| Recraft | Recraft Remove Background | | Quickly removing the background from a single image. | - A dedicated, prompt-free background-removal tool.
| - Only removes backgrounds; it does not provide generative image editing.
|
| Recraft | | | Improving the clarity and detail of a single image. | - A dedicated, prompt-free image-enhancement tool.
| - Focuses on crispness rather than creative generation or detailed controls.
|
| Google | | VideoText to videoFirst frame to videoFirst and last frame to video | High-quality video guided by first and last frames. | - Strong cinematic motion with up to 4K output in the connected workflow.
| - Quality mode is slower and uses more credits than Fast.
|
| Google | | VideoText to videoFirst frame to videoFirst and last frame to video | Rapid first/last-frame video drafts. | - Faster iteration while retaining frame-guided generation.
| - Trades some quality for speed.
|
| Google | | | Video guided by the style or content of reference images. | - Accepts up to three references and supports up to 4K output.
| - References guide the result; they do not provide frame-by-frame control.
|
| xAI | | VideoText to videoImage to video | Flexible short text-to-video or image-to-video concepts. | - Combines text and image modes with durations up to 30 seconds.
| - Current output is limited to 480p or 720p.
|
| xAI | Grok Imagine Video 1.5 Preview | | Short experimental videos anchored to a source image. | - Flexible duration from one to 15 seconds.
| - Preview model; requires an image and is limited to 480p or 720p.
|
| Alibaba | | VideoText to videoFirst frame to videoFirst and last frame to videoVideo continuationReference to video | Text, frame-guided, continuation, and multimodal reference video. | - One model covers five generation modes and up to 1080p.
| - Reference-to-video clips have a shorter maximum duration.
|
| Alibaba | | VideoText to videoFirst frame to videoFirst and last frame to videoReference to video | Text, frame-guided, and image, video, or audio reference workflows. | - Supports flexible clips from two to 30 seconds with output up to 1080p.
| - With reference video, input and output duration combined must stay within 30 seconds.
|
| MiniMax | | VideoText to videoImage to videoReference to video | Reference-rich video with optional image, video, or audio guidance. | - Supports text, frame, and reference-to-video workflows with image, video, or audio references, up to 2K.
| - Input requirements vary by mode and need careful setup.
|
| HappyHorse | | VideoText to videoImage to videoReference to video | Text, image, and multi-reference video creation. | - Combines several generation modes in one model option.
| - Configured for 720p or 1080p, clips from 3 to 15 seconds, and up to nine reference images.
|
| HappyHorse | | | Editing existing video with prompt and image references. | - Supports up to five reference images and 720p or 1080p output.
| - Requires a source video; credit cost varies with source duration and output resolution.
|
| Kuaishou | | VideoText to videoImage to video | Reliable text-to-video and image-to-video generation. | - A focused workflow for common video generation modes.
| - Offers fewer advanced controls than newer Kling variants.
|
| Kuaishou | | VideoText to videoImage to videoFirst and last frame to videoMulti-shot | High-fidelity text and first/last-frame video with native audio. | - Supports Standard, Pro, and 4K tiers with up to five shots.
| - Multi-shot prompts are limited to 15 seconds in total.
|
| Kuaishou | | VideoText to videoImage to video | Faster text or first-frame video iteration. | - Produces three-to-15-second clips at 720p or 1080p.
| - The current integration does not support last-frame guidance.
|
| Kuaishou | | | Long-form 1080p talking avatars driven by a portrait and voice track. | - Produces expressive, multilingual avatar video from one image and one audio file.
| - Requires a prompt; the driving audio must be no longer than five minutes.
|
| Google | | VideoText to videoImage to videoVideo to video | Prompt-led video generation and edits to existing footage. | - Combines creation and editing with reference-image controls.
| - Configured for 720p, 1080p, or 4K at 4, 6, 8, or 10 seconds, with up to seven images and one video reference.
|
| Alibaba | | | Prompt-guided edits to short source videos. | - Supports 720p or 1080p output with an optional reference image.
| - Source clips must be under ten seconds; credit cost varies with duration and resolution.
|
| Topaz Labs | | | Increasing the resolution of existing video. | - A dedicated video-enhancement tool with 1×, 2×, and 4× processing.
| - Requires a source video; credit cost varies with duration and upscale factor.
|
| ByteDance | | | Natural talking-human video generated from one image and one voice track. | - Strong lip sync, facial expression, and lifelike motion at 720p or 1080p.
| - The driving audio must be strictly shorter than 60 seconds.
|
| ByteDance | | VideoText to videoFirst frame to videoFirst and last frame to videoMultimodal reference | Cinematic text, frame-guided, and multimodal-reference video. | - Broad reference support, native audio options, and output up to 4K.
| - The full-quality tier uses more time and credits than Fast or Mini.
|
| ByteDance | | VideoText to videoFirst frame to videoFirst and last frame to videoMultimodal reference | Fast drafts using text, frames, or multimodal references. | - Lower-cost, faster access to the Seedance 2.0 reference modes.
| - Output is limited to 480p or 720p.
|
| ByteDance | | VideoText to videoFirst frame to videoFirst and last frame to videoMultimodal reference | Economical exploration of Seedance multimodal workflows. | - Supports text, frame, image, video, and audio reference modes.
| - Output is limited to 480p or 720p.
|
| ByteDance | | VideoText to videoFirst frame to videoFirst and last frame to videoMultimodal reference | Longer text, frame-guided, and multimodal-reference video workflows. | - Up to 30-second clips, broad multimodal references, and optional native audio.
| - Output is limited to MP4 at 480p, 720p, or 1080p.
|
| Microsoft | | | Voiceover from written scripts. | - Selectable voices and speech speed from 0.5× to 2×.
| - Designed for text-to-speech, not music or sound effects.
|
| Suno | | | Instrumental or vocal music from a prompt. | - Creates complete music ideas with an instrumental option.
| - Exact structure, lyrics, and duration may need prompt iteration.
|
| Suno | | AudioText to sound effect | Prompt-based sound effects and looping ambience. | - Focused sound generation with an optional loop mode.
| - Intended for effects rather than speech or full songs.
|