Skip to main content
I

IntGlobal

Senior DevOps & SRE

Kolkata, West Bengal, IndiaPosted 2 weeks ago

Skills

AWSMySQLKubernetesNode.jsTerraformCI/CDGitDevOpsCollaboration

Job description

Current Tech Stack & Infrastructure  Cloud Infrastructure: AWS (ElastiCache, EKS, RDS Aurora MySQL)  Orchestration: Kubernetes on EKS with Karpenter for node management  Infrastructure as Code: Terraform (multiple repositories, experiencing drift and collaboration challenges)  Monitoring & Alerting: Datadog for monitoring, alerting, incident management, and runbooks  CI/CD: GitHub Actions with Atlantis for infrastructure PRs, Bitbucket Pipelines for Helm deployments, Rundeck for scripted operations  GitOps: ArgoCD for Kubernetes workloads  Data Infrastructure: Transitioning to segmented data stores with ClickHouse, MySQL Aurora, and Pulsar for event streaming  APM: Limited Datadog APM usage due to cost ($45 per host/month, ~10 hosts)  Additional Monitoring: Started Prometheus clusters within EKS for more verbose metrics alongside Datadog Scale & Workload  Handling hundreds of thousands of events per second via API (mix of synchronous and asynchronous processing)  EC2 instances: M5/M6 xlarge instances, scaling between 15-100 instances per environment depending on load  Currently transitioning workloads from EC2 to Kubernetes (K8s workload still smaller than core application)  Real-time workloads requiring minimal downtime (minutes not hours for maintenance windows) Key Operational Challenges  Infrastructure as Code: Terraform has become unmanageable due to drift reconciliation and multi-person collaboration issues  Legacy Environments: Some environments set up entirely manually, never in Terraform, with major operational challenges to migrate while keeping them online  Terraform Migration: Proven migration patterns in test environments, but moving to production challenging due to real-time workload requirements  Alert Management: Receiving too many alerts, need prioritization and structured approach to reduce noise and recategorize/adjust thresholds  Alert Distribution: Historically all alerts went to one tech ops team instead of being distributed to five different dev teams; working to shift left closer to developers  Service Catalog: Still defining service catalog and ownership model  Runbooks: Exist in Datadog but need more polish and structure  Database Scaling: Operational challenges with database scaling being addressed through data store segmentation  Monitoring Costs: Datadog is expensive; moved from CloudWatch ~8 years ago due to cost  Manual Changes: Infrastructure changes often implemented manually first, then imported to Terraform and rolled out to other environments Top 5 Required Skill Sets 1. Strong Terraform and CI/CD process expertise 2. Experience breaking down alerts to identify critical vs. non-critical and handling frequent alarms (Datadog-specific experience preferred) 3. SLA/SLI definition experience (team currently being asked to do this without prior experience) 4. AWS infrastructure expertise, DevOps-heavy background 5. Monitoring and alerting experience (Datadog preferred, though Prometheus/Grafana experience also valuable)

Apply on IntGlobal