Who it is built for
- Engineering teams building production AI agents
- Teams that need dedicated GPU inference capacity
- Organizations with on-premise or VPC deployment requirements
Inference infrastructure for open-weight and custom models
FriendliAI provides API and infrastructure options for running open-weight and custom models in production. Developers can use pay-as-you-go Model APIs, deploy dedicated GPU endpoints for predictable capacity, or run its containerized offering in their own environment for greater deployment control.
The decision
Start with the job, the team, and the constraints. Product fit becomes much clearer when those three line up.
Inside the product
The core product capabilities, grouped around the work they enable.
Serverless Model APIs provide pay-as-you-go access to frontier open-weight models through OpenAI-compatible and Anthropic Messages-style interfaces.
Dedicated Endpoints provide GPU-backed inference with autoscaling, observability, and support for custom and open-source models.
Friendli Container is a Kubernetes-native option for teams that need to run inference with data-protection and governance controls in their own environment.
Sign up for FriendliAI, generate a personal API key, and select an available model. Send inference requests through the API or use the Suite to deploy a dedicated endpoint. Teams with stricter deployment requirements can discuss container or enterprise options with FriendliAI.
Plans and official links
See the entry price, free access options, company details, and direct vendor destinations in one place.
Choosing for a real workflow?
We map the workflow, connect existing systems, choose what to buy, and build what is missing.