
Together AI
Usage BasedThe AI Native Cloud.
Published 2 October 2026
Scores
Popularity3/5
An $8.3B valuation and around $1B in annualized revenue make it one of the largest open-model clouds, and AI engineers know it well; most application developers who only call OpenAI or Anthropic have little reason to.
Learning Curve1/5
An API key and the OpenAI client library with a changed base URL are enough to call any model, and the playground lets newcomers try models before writing code.
Flexibility5/5
Serverless, batch, dedicated endpoints, raw GPU clusters, and fine-tuning cover everything from a quick prototype to training a custom model, across text, image, audio, and embedding workloads.
Performance4/5
Custom kernels from the FlashAttention team and speculative decoding make it one of the faster GPU-based providers for open models, though LPU-based Groq still leads on time to first token.
Portability4/5
Every hosted model is open-weight and the API follows the OpenAI format, so the same model can move to another provider or a self-hosted vLLM server, but fine-tuned weights and dedicated endpoints take effort to migrate.
About Together AI
Together AI is an inference and training cloud built around open-weight models. Instead of running its own frontier model, it hosts more than 200 open models from other labs, including Llama, DeepSeek, Qwen, Kimi, GLM, Mistral, and gpt-oss, along with image, video, audio, embedding, and rerank models, all behind one OpenAI-compatible API. Switching an application from a closed model to an open one is often a matter of changing the base URL and model name.
The platform covers several ways to run a model. Serverless inference bills per token with no setup, and a batch API processes large asynchronous jobs at up to half the price. Provisioned throughput reserves token capacity under an SLA, dedicated endpoints reserve GPUs for a single model with predictable latency, and GPU clusters of H100, H200, B200, and B300 machines can be rented for training or self-managed serving. Fine-tuning jobs, with LoRA or full training and preference methods such as DPO, produce a model that can be deployed on the same platform.
Its engineering edge comes from research: Tri Dao, the author of FlashAttention, is the company's chief scientist, and its inference engine uses custom kernels and speculative decoding to serve open models faster than a stock vLLM deployment. Together also runs a Code Sandbox and Code Interpreter for agents, based on its acquisition of CodeSandbox.
Founded in San Francisco in 2022, Together AI raised an $800M Series C in July 2026 at an $8.3B valuation and reports customers such as Cursor and Decagon. Billing is usage-based from a prepaid balance, with reserved capacity and enterprise contracts for larger workloads.
Key Features
- 200+ open-weight models across chat, code, vision, image, audio, and embeddings
- OpenAI-compatible API and Python and TypeScript SDKs
- Serverless per-token inference with a batch API at up to 50% off
- Provisioned throughput with reserved token capacity and SLAs
- Dedicated endpoints on reserved GPUs for one model
- On-demand, reserved, and preemptible GPU clusters from H100 to B300
- LoRA, full, and DPO fine-tuning with one-click deployment
- Code Sandbox and Code Interpreter for agent code execution
Pros
- The broadest hosted catalog of open-weight models, usually with new releases available within days
- OpenAI-compatible API makes moving an app from a closed model to an open one a small change
- Inference, fine-tuning, and GPU rental on one account, so a model can go from training to serving without changing providers
- Competitive per-token prices on popular open models, with batch processing cutting them further
Cons
- Hundreds of models at different per-token rates make monthly costs hard to predict
- Some reviewers report slow or unhelpful support on billing disputes
- Hosted only: open models avoid lock-in, but the speed optimizations stay on Together's infrastructure
- Serverless rate limits and cold starts on less popular models can push production teams to pricier dedicated endpoints
- No free tier or trial: access starts with a $5 minimum credit purchase
Together AI Pricing
Usage Based- · Per H100 GPU-hour on demand; H200 at $5.99 and B200 at $8.19
- · Reserved terms from $3.19 and preemptible capacity from $1.99 per H100-hour
- · Up to 50% off serverless rates on selected models
- · Results within a 24-hour processing window
- · Reserved token capacity in throughput units (PTUs), billed per PTU-minute
- · Capacity SLAs for steady production traffic
- · Single-tenant GPUs serving one model, billed by the minute per replica
- · B200 at $8.99 per GPU-hour; H200 and B300 quoted by sales
- · Billed per million training tokens, rising with model size
- · LoRA supervised fine-tuning starts at $0.34 per million tokens for small models
- · $4 minimum charge per job
- · $0.0446 per vCPU-hour and $0.0149 per GiB of RAM per hour
- · Code Interpreter sessions at $0.03 per 60 minutes
- · Per-token billing from a prepaid balance, rates set per model
- · Chat models run from $0.09 to $3 per million input tokens
- · No free trial; $5 minimum credit purchase
- · Reserved capacity, private deployments, and custom limits
- · Quoted by sales
Tools Related to Together AI
Integrates with Together AI(1)
LiteLLM has a Together AI provider (the together_ai/ model prefix), so a self-hosted LiteLLM gateway can route requests to Together AI's hosted open models alongside other providers.
Alternatives to Together AI(3)
Together AI and Fireworks AI both serve open-weight models through OpenAI-compatible APIs with fine-tuning and dedicated GPUs. Together AI has the broader catalog and rents raw GPU clusters; Fireworks AI focuses on serving speed for latency-sensitive products. Pick Together AI for breadth and training, Fireworks AI for the fastest responses.
Together AI hosts 200+ open-weight models on GPUs and adds fine-tuning, dedicated endpoints, and GPU clusters; Groq serves a smaller catalog on its own LPU chips with the lowest latency. Pick Together AI for model breadth and training, Groq for raw speed.
Together AI is a focused inference and training cloud with per-token APIs and GPU rental for open models; Hugging Face is the hub where those models are published, with its own inference endpoints on top. Pick Together AI for production serving, Hugging Face for discovery and the widest model selection.