Model, throughput, and context
- Use smaller models for classification, extraction, and normalization; keep evals to catch quality loss.
- Provisioned throughput fits stable high-volume workloads; on-demand fits early or spiky workloads.
- Cross-Region inference profiles can help capacity, but need latency, residency, and compliance review.


