@livingdevopsi
iAccount based inIndia!
About this account
- Account based in
- India
- Connected via
- India App Store
! X says this location may be affected by a proxy or VPN.
Account-level information from X, not a live location or the device used for a specific post.
Founder LivingDevOps | DevOps Lead | Educator | Mentor | Tech Storyteller
Noida, India
Joined February 2023
- Tweets7.4K
- Following602
- Followers22.3K
- Likes6.3K
Pinned Tweet
After 14 years in production DevOps and training 400+ engineers, I just launched my most advanced bootcamp to date.
28 weeks. 56 live classes. DevOps + MLOps + AIOps.
Built for experienced engineers who want to move into MLOps and AIOps without faking it on their resume.
Check it out π
livingdevops.com/courses/28-β¦
Launch offer: up to 30% off if you book by 15th June.
Devops engineers
SRE engineers
Platform engineers
You guys are going to make so much money.
BREAKING: Anthropic is reportedly planning to spend $518,000,000,000.00 on cloud and infrastructure in the coming years.
Akhilesh Mishra retweeted
Agentic AI interview question:
Five agents independently solve the same task. How would you decide whether their results agree?
Akhilesh Mishra retweeted
A few months ago, I was working with a startup. Their stack was simple monolithic app running on EC2.
But their plan wasn't simple. They wanted to break the app into dozens of fine-grained microservices, containerize everything, and run it all on Kubernetes.
I pushed back.
They were a small team with an early-stage product. Dozens of microservices would have meant more deployments, network calls where function calls used to be, distributed debugging, and a lot of moving parts before they even had product-market fit.
Adding Kubernetes on top would have meant a control plane to manage, upgrades to plan, add-ons to maintain, and most likely a dedicated engineer just to keep it all running.
So I suggested a middle path:
β Modular microservices: a few well-bounded services instead of dozens of tiny ones
β Clean module boundaries inside each service, so they can be split further when needed
β Run everything on ECS for now
β Move to Kubernetes when the scale justifies the complexity
Talking a founder out of trendy tech is hard. But recommending the architecture that fits today's problem, not the one that looks good on a pitch deck, is part of the job.
Three hours and about 100 questions later, we agreed.
Then came the next request: "Make sure it scales for every possible traffic spike."
(Every founder is quietly planning for a million daily users. π)
Here's what we built on AWS:
β One ECS cluster running 5 services, connected through ECS Service Connect over HTTP/2
β An ALB with SSL certificates from ACM
β RDS Postgres: a single instance for dev, a Multi-AZ cluster for production. Multiple tables that can be migrated to individual instances in future.
β Multi-environment infrastructure as code
β Separate app and infra repos, with the build pipeline triggering deployments across repos
β CI/CD with GitHub Actions, with security scanning built into CI
β Autoscaling on custom metrics, plus monitoring and alerting
The result is production-ready and scalable, and they don't need a platform team to run it.
When they outgrow it, the boundaries are already in place. Splitting services further and moving to Kubernetes becomes a planned migration, not a rewrite.
I love microservices and Kubernetes. But they aren't day-one defaults.
Start simple. Earn your complexity.
A few months ago, I was working with a startup. Their stack was simple monolithic app running on EC2.
But their plan wasn't simple. They wanted to break the app into dozens of fine-grained microservices, containerize everything, and run it all on Kubernetes.
I pushed back.
They were a small team with an early-stage product. Dozens of microservices would have meant more deployments, network calls where function calls used to be, distributed debugging, and a lot of moving parts before they even had product-market fit.
Adding Kubernetes on top would have meant a control plane to manage, upgrades to plan, add-ons to maintain, and most likely a dedicated engineer just to keep it all running.
So I suggested a middle path:
β Modular microservices: a few well-bounded services instead of dozens of tiny ones
β Clean module boundaries inside each service, so they can be split further when needed
β Run everything on ECS for now
β Move to Kubernetes when the scale justifies the complexity
Talking a founder out of trendy tech is hard. But recommending the architecture that fits today's problem, not the one that looks good on a pitch deck, is part of the job.
Three hours and about 100 questions later, we agreed.
Then came the next request: "Make sure it scales for every possible traffic spike."
(Every founder is quietly planning for a million daily users. π)
Here's what we built on AWS:
β One ECS cluster running 5 services, connected through ECS Service Connect over HTTP/2
β An ALB with SSL certificates from ACM
β RDS Postgres: a single instance for dev, a Multi-AZ cluster for production. Multiple tables that can be migrated to individual instances in future.
β Multi-environment infrastructure as code
β Separate app and infra repos, with the build pipeline triggering deployments across repos
β CI/CD with GitHub Actions, with security scanning built into CI
β Autoscaling on custom metrics, plus monitoring and alerting
The result is production-ready and scalable, and they don't need a platform team to run it.
When they outgrow it, the boundaries are already in place. Splitting services further and moving to Kubernetes becomes a planned migration, not a rewrite.
I love microservices and Kubernetes. But they aren't day-one defaults.
Start simple. Earn your complexity.
Akhilesh Mishra retweeted
That's what I teach in my 9 Week Kubernetes + SRE Bootcamp.
livingdevops.com/courses/9-wβ¦
My most liked tweet
My most retweeted tweet
My most bookmark tweet
My most viewed tweet
Tweet with most replies
Give in this order @grok
Akhilesh Mishra retweeted
We used to ask developers about API contracts, resilience, and application architecture.
We asked DevOps engineers about managing Kubernetes, troubleshooting incidents, and keeping production running.
Now it's flipped.
DevOps engineers get asked about API contracts.
Developers get asked about Kubernetes.
The lines between roles are fading.
Full-stack now means application + cloud + Kubernetes.
DevOps now means infrastructure + architecture + security + observability.
Maybe the industry is moving toward engineers who understand the whole system.
Or maybe we just keep adding requirements to the same job description. π
We used to ask developers about API contracts, resilience, and application architecture.
We asked DevOps engineers about managing Kubernetes, troubleshooting incidents, and keeping production running.
Now it's flipped.
DevOps engineers get asked about API contracts.
Developers get asked about Kubernetes.
The lines between roles are fading.
Full-stack now means application + cloud + Kubernetes.
DevOps now means infrastructure + architecture + security + observability.
Maybe the industry is moving toward engineers who understand the whole system.
Or maybe we just keep adding requirements to the same job description. π
Akhilesh Mishra retweeted
Kubernetes Interview Question That Exposes Real Experience
"Your pods are running. But users can't access the app. What do you do?"
Here's how I'd answer it.
I don't start with the pod logs.
I follow the path a request takes and find where it breaks.
User β DNS β Load Balancer β Ingress β Service β Pod
1. Check if the pod is actually ready
Running is not the same as ready.
If the READY column shows 0/1, the readiness probe is failing.
Kubernetes won't send traffic to that pod.
2. Check the Service endpoints
If the Service has no endpoints, it has nowhere to send traffic.
Usually it's one of two things:
- The selector doesn't match the pod labels
- The pods aren't ready
3. Check the ports
The Service targetPort must match the port the app is listening on.
If they don't match, you get "connection refused" or 502 errors.
4. Check what the app is listening on
This one gets a lot of people.
If the app listens on 127.0.0.1, nothing outside the container can reach it.
It needs to listen on 0.0.0.0.
5. Test from inside the cluster
Spin up a temporary pod and call the Service directly.
If it works, the problem is outside the cluster: Ingress, load balancer, or DNS.
If it fails, the problem is inside: the Service, the pod, or the network.
This one test cuts the problem in half.
6. Check network policies
A deny policy can quietly drop traffic.
Everything looks healthy, but requests never reach the pod.
7. Check the edge
If everything works inside the cluster, look outward:
- Ingress rules: host, path, service name, port
- Load balancer health checks
- Security groups
- DNS records
Interviewers like this answer because
- It shows you understand how traffic flows.
- It shows you debug step by step instead of guessing, and you don't panic and redeploy.
That's how real DevOps and SRE engineers think.
Kubernetes Interview Question That Exposes Real Experience
"Your pods are running. But users can't access the app. What do you do?"
Here's how I'd answer it.
I don't start with the pod logs.
I follow the path a request takes and find where it breaks.
User β DNS β Load Balancer β Ingress β Service β Pod
1. Check if the pod is actually ready
Running is not the same as ready.
If the READY column shows 0/1, the readiness probe is failing.
Kubernetes won't send traffic to that pod.
2. Check the Service endpoints
If the Service has no endpoints, it has nowhere to send traffic.
Usually it's one of two things:
- The selector doesn't match the pod labels
- The pods aren't ready
3. Check the ports
The Service targetPort must match the port the app is listening on.
If they don't match, you get "connection refused" or 502 errors.
4. Check what the app is listening on
This one gets a lot of people.
If the app listens on 127.0.0.1, nothing outside the container can reach it.
It needs to listen on 0.0.0.0.
5. Test from inside the cluster
Spin up a temporary pod and call the Service directly.
If it works, the problem is outside the cluster: Ingress, load balancer, or DNS.
If it fails, the problem is inside: the Service, the pod, or the network.
This one test cuts the problem in half.
6. Check network policies
A deny policy can quietly drop traffic.
Everything looks healthy, but requests never reach the pod.
7. Check the edge
If everything works inside the cluster, look outward:
- Ingress rules: host, path, service name, port
- Load balancer health checks
- Security groups
- DNS records
Interviewers like this answer because
- It shows you understand how traffic flows.
- It shows you debug step by step instead of guessing, and you don't panic and redeploy.
That's how real DevOps and SRE engineers think.
That's what I teach in my 9 Week Kubernetes + SRE Bootcamp.
livingdevops.com/courses/9-wβ¦
Akhilesh Mishra retweeted
Connection pool exhaustion, in plain terms:
It's not "too many users." It's connections being held open longer than they need to be, usually by a slow query or a forgotten transaction.
Add more pool size and you just delay the same crash. Ever chased one of these down?
Akhilesh Mishra retweeted
Kubernetes world is moving away from sidecar patterns.
Let me explain why.
A sidecar is a helper container. It sits next to your app container. Same pod. Same network. Same lifecycle.
People used sidecars for a lot of things.
Service mesh was the big one. Envoy would sit next to your app.
- It handled traffic. It handled retries and timeouts. It handled security between services(MTLS). Your app never knew it was there.
- Logging was another one. A small agent would sit next to your app. It would read the logs. It would ship them somewhere else.
- Metrics scraping worked the same way.
- Secrets injection too. Vault agent would run as a sidecar. It would fetch secrets and write them to a file.
This pattern felt smart. One app. Many helpers. No code changes needed.
But then problems showed up.
Every sidecar is another container.
That means more CPU, more memory.
One sidecar on ten pods is fine.
One sidecar on ten thousand pods is not fine.
Every sidecar can crash.
Every sidecar needs patching.
Every sidecar needs upgrading.
Multiply that by hundreds and thousands.
Startup order became a pain too.
- Sometimes the sidecar was not ready before the app started.
- Sometimes the sidecar died before the app finished its work. Small bugs. Big headaches.
Debugging got harder.
- Now you have two containers to check instead of one.
Cost added up quickly at scale.
So the industry asked a new question.
What if we do not need the sidecar at all?
Istio built something called an ambient mesh that required no sidecar injection.
The proxy moves completely outside the pod.
Cilium went even further by using eBPF (a kernel-level technology).
That means the networking logic lives in the kernel, so no sidecar container, and lightning-fast networking.
OpenTelemetry solved this for metrics and logs.
Instead of one sidecar per pod, you get one collector for many pods(runs as a DaemonSet and deployment).
Even Kubernetes noticed the pain.
In version 1.29, it gave sidecars a proper place in the pod lifecycle, fixing a lot of old bugs and startup issues.
It did not remove the sidecar. It just made it better.
Today we have two paths.
Some teams are removing sidecars completely.
Some teams are keeping sidecars but doing it the right way.
Where do you stand?
Kubernetes world is moving away from sidecar patterns.
Let me explain why.
A sidecar is a helper container. It sits next to your app container. Same pod. Same network. Same lifecycle.
People used sidecars for a lot of things.
Service mesh was the big one. Envoy would sit next to your app.
- It handled traffic. It handled retries and timeouts. It handled security between services(MTLS). Your app never knew it was there.
- Logging was another one. A small agent would sit next to your app. It would read the logs. It would ship them somewhere else.
- Metrics scraping worked the same way.
- Secrets injection too. Vault agent would run as a sidecar. It would fetch secrets and write them to a file.
This pattern felt smart. One app. Many helpers. No code changes needed.
But then problems showed up.
Every sidecar is another container.
That means more CPU, more memory.
One sidecar on ten pods is fine.
One sidecar on ten thousand pods is not fine.
Every sidecar can crash.
Every sidecar needs patching.
Every sidecar needs upgrading.
Multiply that by hundreds and thousands.
Startup order became a pain too.
- Sometimes the sidecar was not ready before the app started.
- Sometimes the sidecar died before the app finished its work. Small bugs. Big headaches.
Debugging got harder.
- Now you have two containers to check instead of one.
Cost added up quickly at scale.
So the industry asked a new question.
What if we do not need the sidecar at all?
Istio built something called an ambient mesh that required no sidecar injection.
The proxy moves completely outside the pod.
Cilium went even further by using eBPF (a kernel-level technology).
That means the networking logic lives in the kernel, so no sidecar container, and lightning-fast networking.
OpenTelemetry solved this for metrics and logs.
Instead of one sidecar per pod, you get one collector for many pods(runs as a DaemonSet and deployment).
Even Kubernetes noticed the pain.
In version 1.29, it gave sidecars a proper place in the pod lifecycle, fixing a lot of old bugs and startup issues.
It did not remove the sidecar. It just made it better.
Today we have two paths.
Some teams are removing sidecars completely.
Some teams are keeping sidecars but doing it the right way.
Where do you stand?
Akhilesh Mishra retweeted
Terraform is great. Until you have to run it across three environments.
You write the code once.
Then you copy it for staging.
Then again for prod.
It feels clean at first. Each environment has its own folder, its own state, its own everything.
A few months later, you fix a bug in dev and forget staging.
Someone patches prod during a late-night hotfix.
A one-line change becomes three PRs.
And nobody can explain why prod behaves differently.
The problem is no longer writing Terraform.
The problem is keeping the same Terraform in sync across environments.
The approach I have found useful is simple:
Keep one codebase. Move the differences into config.
terraform/
βββ modules/
βββ main.tf
βββ variables.tf
βββ vars/
βββ dev.tfvars
βββ dev.tfbackend
βββ prod.tfvars
βββ prod.tfbackend
The code defines the infrastructure.
The .tfvars define the environment.
The .tfbackend defines where the state lives.
Fix something once, and it's fixed everywhere.
The only pain left was the flags:
terraform init -backend-config=vars/prod.tfbackend
terraform plan -var-file=vars/prod.tfvars
Every single time. For every environment.
And one wrong filename away from planning prod with dev values.
That is why I built Toffee.
Instead of remembering flags, you just say which environment:
toffee dev plan
toffee prod apply
It does not replace Terraform.
It does not add a new layer to learn.
It just removes the repetitive part.
That is how I think infrastructure tooling should work.
Less copy-paste.
Fewer flags to remember.
More time building infrastructure.
It's open source; try it and tell me what's missing:
github.com/akhileshmishrabizβ¦
Terraform is great. Until you have to run it across three environments.
You write the code once.
Then you copy it for staging.
Then again for prod.
It feels clean at first. Each environment has its own folder, its own state, its own everything.
A few months later, you fix a bug in dev and forget staging.
Someone patches prod during a late-night hotfix.
A one-line change becomes three PRs.
And nobody can explain why prod behaves differently.
The problem is no longer writing Terraform.
The problem is keeping the same Terraform in sync across environments.
The approach I have found useful is simple:
Keep one codebase. Move the differences into config.
terraform/
βββ modules/
βββ main.tf
βββ variables.tf
βββ vars/
βββ dev.tfvars
βββ dev.tfbackend
βββ prod.tfvars
βββ prod.tfbackend
The code defines the infrastructure.
The .tfvars define the environment.
The .tfbackend defines where the state lives.
Fix something once, and it's fixed everywhere.
The only pain left was the flags:
terraform init -backend-config=vars/prod.tfbackend
terraform plan -var-file=vars/prod.tfvars
Every single time. For every environment.
And one wrong filename away from planning prod with dev values.
That is why I built Toffee.
Instead of remembering flags, you just say which environment:
toffee dev plan
toffee prod apply
It does not replace Terraform.
It does not add a new layer to learn.
It just removes the repetitive part.
That is how I think infrastructure tooling should work.
Less copy-paste.
Fewer flags to remember.
More time building infrastructure.
It's open source; try it and tell me what's missing:
github.com/akhileshmishrabizβ¦
Akhilesh Mishra retweeted
We should start writing our docs in YAML.
AI agents are reading them anyway.
At least weβll save some tokens. π