Job Description
Role Purpose
The purpose of this role is to work with Application teams and developers to facilitate better coordination amongst operations, development and testing functions by automating and streamlining the integration and deployment processes
͏
Do
- Align and focus on continuous integration (CI) and continuous deployment (CD) of technology in applications
- Plan and Execute the DevOps pipeline that supports the application life cycle across the DevOps toolchain â from planning, coding and building, testing, staging, release, configuration and monitoring
- Manage the IT infrastructure as per the requirement of the supported software code
- On-board an application on the DevOps tool and configure it as per the clients need
- Create user access workflows and provide user access as per the defined process
- Build and engineer the DevOps tool as per the customization suggested by the client
- Collaborate with development staff to tackle the coding and scripting needed to connect elements of the code that are required to run the software release with operating systems and production infrastructure
- Leverage and use tools to automate testing & deployment in a Dev-Ops environment
- Provide customer support/ service on the DevOps tools
- Timely support internal & external customers on multiple platforms
- Resolution of the tickets raised on these tools to be addressed & resolved within a specified TAT
- Ensure adequate resolution with customer satisfaction
- Follow escalation matrix/ process as soon as a resolution gets complicated or isnâÂÂt resolved
- Troubleshoot and perform root cause analysis of critical/ repeatable issues
͏
Deliver
| No | Performance Parameter | Measure |
| 1. | Continuous Integration,Deployment & Monitoring | 100% error free on boarding & implementation |
| 2. | CSAT | Timely customer resolution as per TAT Zero escalation |
͏
͏
Key Responsibilities
- Design, build, and manage highly available Kubernetes clusters across hybrid environments (on-premises and cloud platforms such as AWS EKS, Azure AKS).
- Deploy and manage applications manually using tools such as kubectl and Helm, with growing integration of GitOps practices (e.g., ArgoCD).
- Implement and manage observability stacks using Prometheus, Grafana, Loki, and Mimir to monitor infrastructure, applications, and system performance.
- Define, monitor, and improve SLA/SLO/SLI metrics and alerting systems to ensure platform reliability.
- Automate provisioning and configuration of infrastructure using Terraform, Helm, and scripting languages (e.g., Bash, Python).
- Plan, implement, and test backup and disaster recovery (DR) strategies using tools like Velero, Commvault, etc.
- Manage Kubernetes-native networking, storage, and security configurations (Ceph, NFS, Ingress, PodSecurityPolicies, etc.).
- Configure and enforce Kubernetes security best practices using RBAC, OPA/Gatekeeper, NetworkPolicies, and secrets management tools.
- Integrate and operate Kubernetes ecosystem tools such as Karpenter, MicroK8s, Service Meshes, and kubectl plugins.
- Conduct root cause analysis (RCA) and lead resolution efforts for incidents.
- Participate in the on-call rotation for platform availability and incident management.
- Maintain up-to-date documentation, architecture diagrams, runbooks, and SOPs.
- Mentor engineers and advocate for Kubernetes, security, observability, and deployment best practices across teams.
- Continuously stay informed of industry trends in container orchestration, GitOps, security, and cloud-native tooling.
Required Qualifications
- 5+ years of IT/Infrastructure/DevOps experience, with 2+ years in Kubernetes operations in production environments.
- Strong hands-on experience in Kubernetes architecture, cluster operations, and manual application deployment practices.
- Intermediate-level experience in Kubernetes Security, including:
- Cluster hardening, secrets management
- Pod Security Standards (PSS), OPA/Gatekeeper
- Network policies, image scanning, and runtime protections
- Intermediate experience with ArgoCD for GitOps-style Kubernetes deployments.
- Solid proficiency in Linux system administration (Ubuntu, CentOS, RHEL) and troubleshooting.
- Hands-on experience with Kubernetes-native storage (e.g., Ceph, NFS) and persistent volume provisioning.
- Strong familiarity with observability tools: Grafana, Prometheus, Loki, Mimir, etc.
- Proficiency in Infrastructure as Code using Terraform, Helm, and scripting.
- Experience with Velero, Commvault, or similar for backup and DR.
- Experience operating and optimizing cloud-native Kubernetes platforms like EKS, AKS.
- Exposure to tools like Karpenter, MicroK8s, Service Mesh, and Ingress Controllers.
- Familiarity with AI/ML workloads running on Kubernetes is a plus.
- Excellent collaboration, communication, documentation, and incident resolution skills.
Preferred Qualifications
- Kubernetes certifications: CKA, CKAD, or CKS.
- Strong understanding of container security, networking, and distributed system architecture.
- Experience using Portainer for container and Kubernetes management.
- Advanced knowledge of Grafana and other enterprise-grade observability tools.
- Experience managing large-scale Kubernetes clusters (200+ nodes) is highly preferred.
- Prior experience supporting production-grade, high-availability platforms and environments.