
As AI becomes infrastructure
- Source
- LinkedIn Note
- Date
- Link
Claude went down for ~2 hours today. Not catastrophic, but long enough to remind people that AI is becoming core infrastructure before the reliability practices around it have caught up.
Many business platforms are judged against "three nines" expectations (available 99.9% of the time). But at a massive scale, even small gaps create large disruptions downstream. We can expect the model providers to become increasingly more reliable, and I'm sure Anthropic is already working on that, but much of AI reliability will be solved by the application layer becoming more model-agnostic.
AWS has had major outages. Stripe has had outages. Slack and many others have as well. Any tool that becomes core to your business is likely to be a bottleneck at some point.
AI will follow the same pattern, but with a key difference: models are more interchangeable than cloud, payment, or other SaaS platforms usually are. If your payment system is built on Stripe, you can't just switch to another provider on the day of an outage. If you build AI products with the right abstraction layer, though, routing between Claude, OpenAI, Gemini, open models, or specialized models can be automatic.
That makes model flexibility both a product advantage and an infrastructure advantage. This is where tools like OpenRouter, Perplexity, Cursor, and internal model-routing systems become more important. On most days, they're optimization layers; on days like today, they're resilience layers.
Over time, I believe this will become the default. The strongest AI products will be designed flexibly enough to use the best model for the job, fail over when needed, and keep the user experience stable when the model ecosystem underneath them shifts.









