I build AI agents that survive production, and I can show you what breaks. Before that I spent 8 years on backend systems, and most of that time I was also the person who deployed them, watched them, and got woken up when they fell over.
Why the backend years matter for agent work
An agent in production is an operations problem before it is a model problem. The prompt is rarely what fails. What fails is the tool call that times out, the retry that fires twice and charges the customer twice, the retrieval step that hands back a document this user was never allowed to read.
Those are timeout, idempotency and authorization problems. Backend engineering has had answers for them for twenty years. The useful thing I bring to an agent is not a cleverer prompt, it is knowing which of those old answers applies.
How I work
I don't design a system and hand it over. Before anything ships I want three questions answered: how does this fail, what does it cost to run, and can someone else maintain it after me. Cost is one constraint next to latency, retries and blast radius, never the headline.
What I've built
Blip - an open-source uptime monitor that runs entirely on one Cloudflare Worker with D1. No servers, no containers. It watches my clients, my projects and my homelab.
ChronoStash - an open-source backup platform for Postgres, MySQL and MongoDB to any S3-compatible storage. Scheduled, encrypted, with restore that actually works. I built it after learning the hard way what an untested backup is worth.
The homelab - three Proxmox hosts, dual WAN, and open models running on hardware I already owned. Every architecture I write about runs here first, which is why the failures in these posts are ones I watched happen.
The stack
Node.js, TypeScript, NestJS, PostgreSQL, Redis, RabbitMQ, Docker, Kubernetes, Terraform, Cloudflare Workers, Proxmox. On the AI side: RAG with permission-aware retrieval, tenant isolation, and a provider-agnostic LLM layer so OpenAI-compatible APIs, Claude and local Ollama models are all swappable.
Questions I get asked
What does Kazi Shiplu do?
I'm a Lead Engineer in Dhaka, Bangladesh. I design and run AI agent systems in production, and the backend and infrastructure underneath them. 8 years, most of it on call for the systems I built.
Why call an AI agent an operations problem?
A demo agent only has to answer once. A production agent has to answer on the worst day: the tool call times out, the retry fires twice, the retrieval returns a document the user is not allowed to see. Those are latency, idempotency and authorization problems, and they were solved in backend engineering long before anyone called it an agent.
What do you build agents with?
Node.js and TypeScript on NestJS, PostgreSQL and Redis for state, RabbitMQ for anything that must survive a restart, and a provider-agnostic LLM layer so OpenAI-compatible APIs, Claude and local Ollama models are swappable. Deployed on Kubernetes or Cloudflare Workers depending on what the workload needs.
Are your projects open source?
Two are. Blip is an uptime monitoring platform running on one Cloudflare Worker with D1. ChronoStash backs up Postgres, MySQL and MongoDB to S3-compatible storage with restores that have been tested. Both are on GitHub under warlock277.
How do I get in touch?
A LinkedIn DM is the fastest route, email is second. Both are in the footer of every page on this site.
What I write here
I write here about production systems, infrastructure, and what actually breaks. If you're building in that space, find me on LinkedIn. My DMs are open.
