Recode Solutions
AI/ML Architect
Skills
Job description
Job Summary We are seeking an experienced AI/ML Architect to lead the design and development of scalable, real-time AI systems. You will work closely with product, data, and engineering teams to architect end-to-end solutions — from model development and deployment to system integration and production monitoring. Key Responsibilities · Design and architect AI/ML systems that are scalable, low-latency, and production-ready · Lead development of real-time inference pipelines for use cases like voice, vision, or NLP · Select and integrate appropriate tools, frameworks, and infrastructure (e.g., Kubernetes, Kafka, TensorFlow, PyTorch, ONNX, Triton, VLLM etc.) · Collaborate with data scientists and ML engineers to productionize models · Ensure reliability, observability, and performance of deployed systems · Conduct architecture reviews, POCs, and system optimizations · Mentor engineers and help set best practices for ML lifecycle (MLOps) Requirements · 6+ years of experience building and deploying ML systems in production · Proven expertise in real-time, low-latency system design (e.g., streaming inference, event-driven pipelines) · Strong understanding of scalable architectures — microservices, message queues, distributed training/inference · Proficient in Python and popular ML/DL frameworks (scikit-learn, TensorFlow, PyTorch) · Hands-on experience with LLM inference optimization using frameworks like vLLM, TensorRT-LLM, and SGLang · Familiarity with vector databases, embedding-based retrieval, and RAG pipelines · Experience with containerized environments (Docker, Kubernetes) and managing multi-container applications · Working knowledge of cloud platforms (AWS, GCP, or Azure) and CI/CD practices for ML workflows · Exposure to edge deployments and model compression/optimization techniques · Strong foundation in software engineering principles and system design Nice to Haves · Experience in Linux (Ubuntu) · Terminal/Bash Scripting