Compact GPT model for low-latency assistance and high-volume workloads
Compact GPT model for low-latency assistance and high-volume workloads. Supports a context window of 128,000 tokens. Supports tool calling. Input priced at $10 per million tokens. Weights are not publicly released; access is via the provider API.
OpenAI
api
paid
No benchmark results have been added yet.