Input is processed in parallel — the prefill phase —, so its cost is mostly about compute. Output is generated one token at a time — the decode phase —, limited by how much memory has to move for each one: that's why it scales worse, and why the price list always makes output more expensive. That's where the three dials below come from: a cache hit reuses a prefill already done (which is why it costs ~10%), and reasoning tokens are pure decode — they are billed as output because, technically, that is exactly what they are.
Costs
Quantization