# Baseten

AI inference platform offering token-priced model APIs and dedicated GPU deployments of customers' own models, billed per minute and able to scale to zero.


Source: https://www.findhost.app/baseten/
Updated: 2026-09-24
Listed since: 2026-09-24

## Record

### Links

- home: https://www.baseten.co
- pricing: https://www.baseten.co/pricing/
- status: https://status.baseten.co
- terms: https://www.baseten.co/terms-and-conditions/
- docs: https://docs.baseten.co/overview

### Identity

- founded: unknown
- headcount: unknown
- referringSubnets: unknown
- favorite: unknown
- ownership: unknown
- parent: unknown
- status: active
- wikidata: unknown

### Classification

- category: serverless, gpu
- whoManagesOs: self-managed
- infraContract: unknown
- runsOn: unknown
- cdnFrom: unknown
- useCases: ai-app, api
- audience: unknown

### Tech stack

- panels: unknown
- runtimes: python, docker
- software: unknown
- specialisation: unknown
- sshAccess: unknown
- managedDatabases: unknown
- gpuCapacity: inference, serverless, model-api
- persistentStorage: unknown
- backupsIncluded: unknown

### Deployment

- deployMethods: docker-image

### Pricing

- pricingModel: usage-based
- priceFrom: xl
- entryPrice: {"amount":0.63,"currency":"USD","period":"hour"}
- priceTo: 3xl
- currencies: USD
- billingPeriods: monthly
- billingTiming: arrears
- exitWithin: unknown
- moneyBack: unknown
- renewalMultiple: unknown
- freeTier: trial
- contractMinimum: unknown
- paymentMethods: unknown

### Regions and law

- regions: US
- hqCountry: US
- gdprDpa: unknown

### Other features

- energyClaim: unknown
- certifications: unknown
- greenWebId: unknown
- domainRegistration: unknown
- dnsHosting: unknown
- emailHosting: unknown
- cdnIncluded: unknown
- testDomain: unknown
- collaboration: unknown
- staging: unknown

### Support

- supportChannels: unknown
- supportHours: unknown
- supportTiering: unknown
- sla: unknown

### Automation

- mcpServer: official
- cliTool: official
- apiAvailable: public
- iacSupport: unknown

### Sources

- category: https://docs.baseten.co/overview (read 2026-09-24)
- hqCountry: https://www.baseten.co/terms-and-conditions/ (read 2026-09-24)
- whoManagesOs: https://docs.baseten.co/development/model/custom-server (read 2026-09-24)
- useCases: https://docs.baseten.co/overview (read 2026-09-24)
- runtimes: https://docs.baseten.co/development/model/model-class (read 2026-09-24)
- runtimes: https://docs.baseten.co/development/model/custom-server (read 2026-09-24)
- deployMethods: https://docs.baseten.co/development/model/custom-server (read 2026-09-24)
- gpuCapacity: https://docs.baseten.co/inference/model-apis/overview (read 2026-09-24)
- gpuCapacity: https://docs.baseten.co/development/model/build-your-first-model (read 2026-09-24)
- gpuCapacity: https://docs.baseten.co/deployment/manage/scaling (read 2026-09-24)
- pricingModel: https://www.baseten.co/pricing/ (read 2026-09-24)
- priceFrom: https://www.baseten.co/pricing/ (read 2026-09-24)
- priceTo: https://www.baseten.co/pricing/ (read 2026-09-24)
- entryPrice: https://www.baseten.co/pricing/ (read 2026-09-24)
- currencies: https://www.baseten.co/pricing/ (read 2026-09-24)
- billingPeriods: https://docs.baseten.co/organization/billing (read 2026-09-24)
- billingTiming: https://docs.baseten.co/organization/billing (read 2026-09-24)
- freeTier: https://docs.baseten.co/organization/billing (read 2026-09-24)
- regions: https://docs.baseten.co/deployment/regional-deployments (read 2026-09-24)
- apiAvailable: https://docs.baseten.co/deployment/manage/overview (read 2026-09-24)
- cliTool: https://docs.baseten.co/deployment/manage/overview (read 2026-09-24)
- mcpServer: https://docs.baseten.co/agent-setup (read 2026-09-24)

## Notes

Baseten sells AI inference in two shapes. Model APIs are OpenAI- and Anthropic-compatible endpoints over a catalog of open models the company runs, priced per million tokens. Dedicated deployments run the customer's own model, packaged with the Truss framework from Python code and a config file, or as a custom Docker container running an inference server such as vLLM or SGLang, on GPUs billed per minute. Deployments autoscale, and the [scaling docs](https://docs.baseten.co/deployment/manage/scaling) show how to set the minimum to zero so an idle model releases its replicas. Training jobs on managed GPUs sit beside inference, and checkpoints can be deployed directly.

The self-serve plan has no monthly fee and charges usage; higher plans and self-hosted deployments are sold through sales. The contracting entity is Baseten Labs, Inc., in San Francisco.

## Worth knowing

New workspaces receive credits, and the [billing docs](https://docs.baseten.co/organization/billing) say models are deactivated when those run out and no payment method is on file. Invoices are issued when usage passes a threshold or at the end of the calendar month, whichever comes first. Deployments can be pinned to a United States or a European Union region, and regional selection requires a verified organization.

---

Data licensed CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). Attributes are recorded, never scored; absent means unknown.
Credit, in full: FindHost, findhost.app, CC BY 4.0
