πŸ” Search
Sign in to post
Show HN: I built a tool showing how AI providers (should) throttle their modelshttps://throttle.staffinganalytics.io

OP here: this project was born out of the frustration/paranoia that AI providers are throttling their models when their server load is too high. So, I set out to model and study the problem mathematically to understand what was happening, what I found was quite surprising. The idea seems natural: as the data center demand increases momentarily through the day, throttling their models (either using quantized versions, reducing the context window or lowering the tier of the model to a smaller one) seems appealing as the replacement model in principle uses less electricity. The problem is that t…

β†—

0trust.social media

Loading your media...

Pick a GIF β€” Giphy

Loading GIFs...