NPCI International Payments Limited
Senior Associate ART Engineer
Skills
Job description
The opportunity NPCI is looking for a highly skilled DevOps Engineer specializing in automation, Kubernetes, and system reliability engineering, with a strong focus on AI-driven testing and resilience engineering. In this role, you will architect and execute automated reliability testing frameworks, simulate real-world failure scenarios, and ensure high availability, fault tolerance, and scalability of mission-critical payment systems operating at national scale. This is a high-impact role where you'll work at the intersection of: DevOps + SRE + Chaos Engineering + AI-driven testing Ensuring platforms are failure-ready, not just failure-resistant Job details Job Title: DevOps Engineer – Reliability & Automation Division: Quality Control & Monitoring Systems Experience: 3–7 Years Employment Type: Full-time Location: Hyderabad Education: Bachelor’s degree in Computer Science / Engineering or related field Preferred Certifications: CKA / CKAD / Cloud DevOps (AWS/Azure/GCP) Work Mode: 5 Days Work From Office (WFO) Key Responsibilities DevOps & Platform Engineering • Design, implement, and manage CI/CD pipelines for application delivery and infrastructure provisioning • Containerize applications using Docker and orchestrate via Kubernetes • Deploy, manage, and optimize production-grade Kubernetes clusters Reliability & Resilience Engineering • Design and implement automated resilience and reliability testing frameworks • Develop and execute failure simulation strategies (node failures, latency, network partitioning, etc.) • Build and run fault injection mechanisms for controlled chaos testing • Identify system weaknesses and improve fault tolerance & recovery strategies Monitoring, Observability & Incident Readiness • Implement monitoring, logging, and alerting solutions (Prometheus, Grafana, ELK/EFK, Open Telemetry) • Define and track SLIs, SLOs, and error budgets • Enhance observability for distributed systems Infrastructure, Networking & Security • Manage Linux systems and troubleshoot performance and reliability issues • Configure networking components including: ➤ Load balancing ➤ DNS ➤ Firewalls & VPN • Ensure security, compliance, and best practices across infrastructure Collaboration & Productivity • Work closely with SRE, DevOps, QA, and development teams • Improve developer productivity through automation and self-service platforms • Contribute to platform reliability culture and engineering excellence Requirements Required Technical Skills Core DevOps Stack Strong hands-on experience with: Docker & Kubernetes (Administration + Automation) CI/CD tools: Jenkins, GitLab CI, GitHub Actions, ArgoCD, Flux Experience in Infrastructure as Code (IaC): Terraform, Ansible, Helm, Pulumi Systems & Networking Expertise Strong Linux administration skills (Ubuntu, CentOS, RHEL) Deep understanding of networking: TCP/IP, DNS, routing Load balancing, firewalls, VPN Observability & Monitoring Experience with: Prometheus Grafana ELK/EFK stack Open Telemetry Programming & Automation Proficiency in: Bash / Shell Python / Go Strong scripting for automation and testing frameworks Systems Thinking Understanding of: Distributed systems behavior Failure modes & recovery patterns Storage systems (Ceph, physical storage) Network stack (Overlay networking) Good to Have Skills and Experience Required Advanced Reliability Engineering • Exposure to Chaos Engineering frameworks (Chaos Mesh, Litmus, Gremlin) • Understanding of SRE practices and production reliability AI-Driven Testing • Experience with AI-powered testing or automation tools • Familiarity with intelligent fault detection & predictive failure analysis Distributed Systems Expertise Understanding of: • System behavior under large-scale failures • High-throughput, low-latency architectures Security & Advanced Platform Knowledge of: • Kubernetes security (RBAC, Network Policies) • GitOps methodologies • Advanced networking (CNI plugins, eBPF)