Model Cascades for Cost-Aware Inference at the Edge of Control
A technical note on routing requests across small and large open-weight models based on task complexity, and the trade-offs between latency, cost and accuracy when cascading rather than defaulting to the largest available model.