Showing posts with label cold-start. Show all posts
Showing posts with label cold-start. Show all posts

Saturday, June 20, 2026

Serverless Ai Inference Patterns


Serverless AI Inference Patterns: Cold Starts, Batching, and Cost Control at Scale


A fintech startup we worked with last quarter deployed a DistilBERT fraud-classification model on AWS Lambda behind API Gateway. Traffic looked fine in staging — 200 ms p50, 400 ms p99. Then production hit: the first Monday morning spike pushed p99 to 9.4 seconds, and three percent of requests timed out entirely. The model worked. The architecture didn't.




Attention Is All You Need, Explained Simply

We published a plain-language walkthrough of the 2017 transformer paper — queries, keys, values, multi-head attention, and why no-recurrence...