AgentHub

Agent Directory

Browse 9,293 AI agents

Status

112 of 245

Auto Router
Live

Your prompt will be processed by a meta-model and routed to one of dozens of models (see below), optimizing for the bes…

api
Design
No ratings
OpenAI: GPT-4 Turbo
Live

The latest GPT-4 Turbo model with vision capabilities. Vision requests can now use JSON mode and function calling. Tra…

api
Design
No ratings
OpenAI: GPT-4o (2024-05-13)
Live

GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It mainta…

api
Design
No ratings
OpenAI: GPT-4o-mini (2024-07-18)
Live

GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs w…

api
Design
No ratings
Amazon: Nova Lite 1.0
Live

Amazon Nova Lite 1.0 is a very low-cost multimodal model from Amazon that focused on fast processing of image, video, a…

api
Design
No ratings
Qwen: Qwen2.5 VL 72B Instruct
Live

Qwen2.5-VL is proficient in recognizing common objects such as flowers, birds, fish, and insects. It is also highly cap…

api
Design
No ratings
Google: Gemma 3 27B
Live

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 1…

api
ResearchDesign
No ratings
Google: Gemma 3 12B
Live

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 1…

api
ResearchDesign
No ratings
Baidu: ERNIE 4.5 VL 424B A47B
Live

ERNIE-4.5-VL-424B-A47B is a multimodal Mixture-of-Experts (MoE) model from Baidu’s ERNIE 4.5 series, featuring 424B tot…

api
Design
No ratings
ByteDance: UI-TARS 7B
Live

UI-TARS-1.5 is a multimodal vision-language agent optimized for GUI-based environments, including desktop interfaces, w…

api
Design
No ratings
Qwen: Qwen3 VL 235B A22B Instruct
Live

Qwen3-VL-235B-A22B Instruct is an open-weight multimodal model that unifies strong text generation with visual understa…

api
Design
No ratings
Qwen: Qwen3 VL 235B A22B Thinking
Live

Qwen3-VL-235B-A22B Thinking is a multimodal model that unifies strong text generation with visual understanding across …

api
ResearchDesign
No ratings