Vertiv
Application Development & Support Specialist
Skills
Job description
Position Overview
The AI/ML Engineer is a hands-on technical specialist responsible for designing, building, and operationalizing AI and machine learning capabilities within the MDM platform and its surrounding ecosystem. This role focuses on applying AI/ML to core MDM problems — entity matching, deduplication, data quality scoring, anomaly detection, and intelligent automation of data stewardship workflows.
This is not a research role — the AI/ML Engineer builds production-grade models and integrates them into the MDM platform's operational pipelines. The role operates under the "you build it, you support it" model, owning AI/ML features from experimentation through production deployment and ongoing monitoring.
The AI/ML Engineer works closely with the MDM Architect (who drives overall technical direction), the Engineering Manager (who drives delivery), and MDM/Integration developers (who build the platform and pipelines that AI/ML models plug into).
Key Responsibilities AI/ML Model Development & Integration
Design, build, and train ML models for entity matching, deduplication, and record linkage across customer, supplier, contact, and item domains
Develop probabilistic and deterministic matching algorithms that improve match accuracy over traditional rule-based approaches
Build data quality scoring models that assess completeness, accuracy, consistency, and timeliness of master data records
Develop anomaly detection models to identify data quality issues, unusual patterns, and potential duplicates in real-time data flows
Design and implement intelligent survivorship logic using ML to determine optimal golden record attribute values from multiple sources
Build NLP/text processing capabilities for entity name standardization, address parsing, and fuzzy matching
Integrate AI/ML models into MDM platform workflows — matching, merging, stewardship routing, and exception handling
Develop automated data classification and categorization models for incoming records
Build recommendation engines for data stewards — suggest merge candidates, flag potential false positives, prioritize review queues
Experiment with graph-based approaches for relationship discovery and network analysis across MDM entities
MLOps & Production Operations
Deploy ML models to production environments with proper versioning, monitoring, and rollback capabilities
Build and maintain ML pipelines for model training, validation, and deployment (e.g., MLflow, Kubeflow, SageMaker, Azure ML)
Implement model monitoring — track prediction accuracy, data drift, concept drift, and model degradation over time
Design A/B testing frameworks to compare model performance against rule-based baselines
Build automated retraining pipelines triggered by performance degradation or data distribution changes
Manage feature stores and feature engineering pipelines for MDM-specific features
Optimize model inference performance for real-time matching scenarios (latency, throughput)
Maintain model documentation including training data, hyperparameters, performance metrics, and decision thresholds
AI Platform & Agentic Workflows
Design and build agentic AI workflows integrated with MDM processes (e.g., using AWS Bedrock, LangChain, or similar frameworks)
Develop LLM-powered capabilities for data enrichment, entity extraction, and intelligent data validation
Build AI-assisted stewardship tools that reduce manual review effort through intelligent automation
Implement RAG (Retrieval-Augmented Generation) patterns for contextual data quality recommendations
Evaluate and integrate foundation models and LLMs for MDM-specific use cases
Design prompt engineering strategies and guardrails for production LLM integrations
Build conversational interfaces for data stewards to query and interact with MDM data using natural language
Data Engineering for AI/ML
Design and build feature engineering pipelines that extract ML-ready features from MDM, ERP, CRM, and data lake sources
Build training data pipelines — extract, label, and version training datasets from production MDM data
Implement data preprocessing, cleansing, and normalization pipelines specific to ML model inputs
Collaborate with data lake and integration teams to ensure AI/ML pipelines have access to required data sources
Design and maintain data schemas for ML feature stores and model input/output contracts
Production Support & Troubleshooting
Own production support for AI/ML features — monitor model performance, investigate prediction failures, and resolve issues within SLAs
Debug model prediction errors — trace through feature extraction, model inference, and post-processing to isolate root cause
Analyze model logs and prediction outputs to identify systematic errors or bias
Collaborate with MDM developers when AI/ML model outputs cause downstream data quality issues
Perform root cause analysis when match/merge accuracy degrades and implement corrective actions
Maintain runbooks for AI/ML model operations, retraining procedures, and incident response
Testing & Quality
Design and execute model evaluation frameworks — precision, recall, F1, AUC for matching models
Build automated test suites for model validation including edge cases, boundary conditions, and adversarial inputs
Conduct A/B testing and champion/challenger experiments to validate model improvements
Perform bias and fairness testing across different data segments and domains
Participate in code reviews for ML code and provide feedback on data pipeline quality
Maintain regression test datasets to ensure model updates don't degrade performance on known scenarios
Security & Compliance
Working knowledge of data security, information security practices, and SOX compliance for AI/ML model changes in production
Ensure PII/sensitive data handling compliance in model training and inference pipelines
Implement model explainability and audit trails for compliance-sensitive matching decisions
Required Qualifications AI/ML Expertise
10 - 12 years hands-on experience building and deploying ML models to production (not just research/experimentation)
4+ years experience with entity matching, record linkage, or deduplication using ML approaches (e.g., probabilistic matching, deep learning for entity resolution)
4+ years experience with Python ML ecosystem — scikit-learn, pandas, NumPy, TensorFlow or PyTorch
3+ years experience with NLP/text processing for entity name matching, address parsing, fuzzy matching (e.g., spaCy, NLTK, Hugging Face Transformers)
3+ years experience with MLOps tools and practices — model versioning, deployment, monitoring (e.g., MLflow, Kubeflow, SageMaker, Azure ML, Vertex AI)
2+ years experience with LLMs and generative AI — prompt engineering, RAG patterns, agentic workflows (e.g., AWS Bedrock, OpenAI API, LangChain, Claude)
Experience with graph-based ML or network analysis techniques (e.g., Neo4j, GraphSAGE, node2vec)
Experience with feature engineering and feature store management (e.g., Feast, Tecton, SageMaker Feature Store)
Data Engineering & Platform Skills
5+ years advanced SQL experience — complex queries, performance tuning, data analysis (e.g., Oracle, SQL Server, PostgreSQL, Snowflake)
4+ years hands-on coding in Python — production-quality code, not just notebooks (including testing, error handling, logging)
3+ years experience building data pipelines for ML — feature extraction, training data preparation, model serving (e.g., Apache Spark, Airflow, dbt)
2+ years experience with cloud ML platforms and services (e.g., AWS SageMaker, Azure ML, GCP Vertex AI)
2+ years experience with containerization for model deployment (e.g., Docker, Kubernetes)
Experience with CI/CD for ML models — automated testing, deployment, and rollback (e.g., Jenkins, GitLab CI, GitHub Actions)
Experience with log analysis and monitoring tools for model observability (e.g., Splunk, ELK, CloudWatch, Prometheus/Grafana)
MDM & Data Quality Domain Knowledge
3+ years experience working with MDM platforms or data quality systems (e.g., Informatica MDM, Reltio, Informatica DQ, Ataccama)
Understanding of MDM concepts: match/merge, survivorship, golden record, hierarchy management, stewardship workflows
Experience with data quality dimensions — completeness, accuracy, consistency, timeliness, uniqueness
Experience with incident management using ticketing systems (e.g., ServiceNow, Jira)
Working knowledge of SDLC: development, testing, CI/CD, change management, release management
Preferred Qualifications
Manufacturing industry experience — customer, supplier, item master data
Experience applying ML to multi-domain MDM (customer, supplier, contact, item)
Experience with Informatica MDM or Reltio platform internals and extensibility
Experience with graph databases for relationship modeling (e.g., Neo4j, Amazon Neptune)
Experience with real-time ML inference at scale (low-latency model serving)
Experience with data labeling and annotation workflows for training data creation
Experience with model explainability frameworks (e.g., SHAP, LIME)
Publications or patents in entity resolution, record linkage, or data quality
Familiarity with event-driven architecture and streaming ML (e.g., Kafka + ML inference)
Agile/Scrum delivery experience