// pricing

You pay for the GPU your model needs.

Every model runs on its own GPU. A bigger model needs a bigger GPU, and a bigger GPU costs more per hour. Pilot picks the smallest one that fits, and your per-token price follows from that GPU’s hourly cost. No subscription, no minimum plan.

// free pre-flight

Try any model for free

Every launch is a free Pre-flight today: the model stays up for 15 minutes or 25 requests, whichever comes first, then it is torn down. Up to 2 Pre-flights per account. Paid, always-on launches are opening soon.

// dedicated launches

Price by model size

A dedicated launch is billed from your wallet: a one-time start fee, then per-token prices set by the GPU. A model that fits a small GPU costs far less per token than one that needs a large GPU.

MODEL SIZEGPU MEMORYTYPICAL GPUSGPU $/HRINPUT / 1M TOKENSOUTPUT / 1M TOKENSSTART FEE
up to ~4B16 GBRTX A4000 / A4500 / 4000 Ada$0.58$0.30$1.21$0.021
~4B – 9B24 GBL4 / RTX A5000 / RTX 4090$0.69 – $1.10$0.36 – $0.57$1.44 – $2.29$0.025 – $0.040
~9B – 10B32 GBRTX 5090 / RTX PRO 4500$1.15 – $1.58$0.60 – $0.82$2.40 – $3.29$0.042 – $0.057
10B, long context48 GBA40 / L40S / RTX 6000 Ada$1.22 – $1.75$0.64 – $0.91$2.54 – $3.65$0.044 – $0.063

// estimates for 16-bit weights; quantized models can fit a smaller GPU. Pilot makes the final pick at launch time.

// how it is calculated

From GPU hour to token price

  • Output tokens = GPU $/hr ÷ (200 tokens/s × 3600) × 1,000,000 × 1.5.
  • Input tokens cost a quarter of output tokens, since reading a prompt is much cheaper than generating a reply.
  • Start fee is what the GPU costs while your model cold-starts (about 130 seconds of its hourly price). It is charged once, when the launch succeeds.
  • GPUs scale to zero when idle, so you are not billed for hours nobody calls the model. A model unused for 7 days is deleted.

Launch your first model free

Sign in with your litepod account. No card needed for a Pre-flight.

get started →