How the current generation of text-to-speech models actually differ across latency, naturalness, and cost, and which tier fits which use case.
How the current generation of text-to-speech models actually differ across latency, naturalness, and cost, and which tier fits which use case.
Platforms
Launches
Platforms
A two-stage rollout, frontier-tier pricing, and a benchmark chart with at least one number that should make you suspicious.
Flash tiers don't win benchmark charts. They win invoices — and this one lands thirteen times under the frontier rate.
No preview stage, no waitlist — Anthropic's new top tier went live everywhere on day one. Here's what to re-test before migrating.
Context bills every call. An index bills once. That arithmetic has outlived every context window expansion so far.
A quantized 27B model fits in 24 GB and handles most routine work. Throughput, not capability, is what actually limits it.
Frontier to budget spans three orders of magnitude. Picking the right rung matters more than picking the right vendor.
Uneven attention, rate limits below the advertised window, and linear cost on every call. Useful — and frequently misapplied.
Free weights are not free inference. The cost case lives on a utilization curve, and residency is a better reason anyway.
Seven things worth extracting from a launch post, in the order that matters if you actually have to ship on the thing.