Agivant
MLOps Engineer
Skills
Job description
We are seeking a highly skilled MLOps Engineer to design, implement, and manage the infrastructure and deployment pipelines for Machine Learning (ML) systems. The ideal candidate will have deep expertise in CI/CD, container orchestration, cloud platforms, and observability tools to ensure performance, scalability, and reliability of AI-driven workflows. Key Responsibilities ML Infrastructure & Deployment: Implement best practices in MLOps, including automated model deployment, versioning, and monitoring using tools such as Kubernetes, Docker, MLflow, Kubeflow, and Azure ML. CI/CD Pipeline Development: Design, implement, and maintain CI/CD pipelines for ML systems using tools such as Jenkins, GitHub Actions, or GitLab CI. Performance & Reliability Monitoring: Continuously monitor and optimize the performance, scalability, and reliability of ML systems. Infrastructure Scaling: Build and scale infrastructure to support AI workflows across development, staging, and production environments. Observability & Monitoring: Deploy and manage observability tools (Prometheus, Grafana, ELK Stack, Azure Monitor) for real-time system health tracking and alerting. Vector Database Management: Set up and maintain infrastructure for vector databases (e.g., Pinecone, Weaviate, Milvus) to support AI-driven applications. Requirements Required Skills & Qualifications Experience: Minimum 3+ years in DevOps/Infrastructure and 2+ years in MLOps. ML Tools & Frameworks: Experience with MLflow, Kubeflow, and model lifecycle management. CI/CD Expertise: Hands-on experience with Jenkins, GitHub Actions, GitLab CI, Azure DevOps. Programming & Scripting: Proficiency in Python, Bash/Shell scripting, and YAML for automation and configuration. Containerization & Orchestration: Strong experience with Docker, Kubernetes, Helm. Infrastructure as Code (IaC): Proficiency in Terraform, Ansible, or CloudFormation. Cloud Platforms: Expertise in AWS, GCP, or Azure (preferred: Azure ML). Monitoring & Observability: Familiarity with Prometheus, Grafana, ELK Stack, Azure Monitor.