> coding agent spend is quietly hockey-sticking across every eng org right now
> the instinct is to cap usage or downgrade models, but that just trades cost for quality
> Not Diamond takes a smarter path: routes model + reasoning effort at every single turn, optimizing for the full session instead >of the next request
> 20%+ lower inference cost, same frontier-level output
> this is the infra layer nobody's thinking about yet but everyone will need
> the instinct is to cap usage or downgrade models, but that just trades cost for quality
> Not Diamond takes a smarter path: routes model + reasoning effort at every single turn, optimizing for the full session instead >of the next request
> 20%+ lower inference cost, same frontier-level output
> this is the infra layer nobody's thinking about yet but everyone will need
1
8
78
11.6K









