How is a GPU request billed?
By its input tokens, at the model's price per million, and at least 256 tokens a request: a request with eight questions is still one request. Output tokens are free, because System One models read and decide without writing text. A request that fails costs nothing.
What does a model cost?
Each model's price is what its GPU costs per million input tokens, measured by an audit on our GPUs, plus our margin; an audit never prices a model above $0.04. Models on our CPU servers answer free within your plan's decisions.
What happens when my plan’s tokens run out?
GPU answers carry on, paid from prepaid credit at each model’s price: yours, or for an organization’s private models its payer’s. Without credit, models on our CPU servers still answer within your decisions. The included tokens renew on the 1st of each month (UTC).
Can I buy a plan today?
Not online yet: buying online is coming soon. Until then our team grants plans by hand; write to [email protected] and say which plan you want. Credit packs can be bought on Settings → Billing, and every verified account gets $5 of free credit.
What is a Dedicated GPU?
A GPU with 24 GB or 48 GB of memory that our team sets up for one organization’s requests: once it is set up, nobody else’s requests queue in front of yours, and until then they run on our shared GPUs. Nothing is charged per token from credit; it is billed monthly, by agreement. Working alone? We set it up on an organization of your own.