Agent Directory
Browse 9,293 AI agents
1–12 of 245
Your prompt will be processed by a meta-model and routed to one of dozens of models (see below), optimizing for the bes…
The latest GPT-4 Turbo model with vision capabilities. Vision requests can now use JSON mode and function calling. Tra…
GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It mainta…
GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs w…
Amazon Nova Lite 1.0 is a very low-cost multimodal model from Amazon that focused on fast processing of image, video, a…
Qwen2.5-VL is proficient in recognizing common objects such as flowers, birds, fish, and insects. It is also highly cap…
Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 1…
Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 1…
ERNIE-4.5-VL-424B-A47B is a multimodal Mixture-of-Experts (MoE) model from Baidu’s ERNIE 4.5 series, featuring 424B tot…
UI-TARS-1.5 is a multimodal vision-language agent optimized for GUI-based environments, including desktop interfaces, w…
Qwen3-VL-235B-A22B Instruct is an open-weight multimodal model that unifies strong text generation with visual understa…
Qwen3-VL-235B-A22B Thinking is a multimodal model that unifies strong text generation with visual understanding across …