Multimodal embedding model mapping text, images, video, audio, and PDFs into a unified embedding space
Multimodal embedding model mapping text, images, video, audio, and PDFs into a unified embedding space. Supports a context window of 8,192 tokens. Input priced at $0.2 per million tokens. Weights are not publicly released; access is via the provider API.
api
paid
No benchmark results have been added yet.