SAI Group Ltd
AI Engineer — LLM / VLM
Skills
Job description
Role Overview
We are looking for an AI Engineer specializing in Large Language Models (LLMs) and Vision-Language Models (VLMs) to design, develop, and deploy production-grade AI solutions. The ideal candidate should have strong experience with LLM/VLM architectures, prompt engineering, RAG, fine-tuning, multimodal AI, and model serving.
Key Responsibilities
Develop and deploy AI applications using LLMs and VLMs .
Build RAG pipelines involving document ingestion, chunking, embeddings, retrieval, reranking, and generation.
Work with models such as GPT, Claude, Gemini, Llama, Mistral, Qwen, and multimodal/VLM models .
Develop multimodal solutions involving text, images, PDFs, charts, tables, and documents .
Perform prompt engineering, supervised fine-tuning, LoRA/QLoRA, and model evaluation .
Build AI agents and tool-calling workflows where appropriate.
Optimize inference for latency, throughput, memory, and cost .
Develop APIs and production services using Python, FastAPI, Docker, and cloud platforms .
Implement evaluation frameworks to measure accuracy, hallucination, relevance, latency, and safety .
Collaborate with ML engineers, software engineers, and product teams to take prototypes into production.
Required Skills
Strong Python programming and software-engineering fundamentals.
Hands-on experience with LLMs and/or VLMs .
Strong understanding of Transformers, attention mechanisms, tokenization, embeddings, and inference .
Experience with PyTorch and Hugging Face Transformers.
Experience building RAG systems and vector-search solutions.
Knowledge of prompt engineering and LLM evaluation .
Experience with APIs, REST services, Git, Docker, and CI/CD.
Familiarity with vector databases such as FAISS, Milvus, Pinecone, Weaviate, or pgvector .
Understanding of cloud AI infrastructure, preferably AWS/Azure/GCP.
VLM / Computer Vision Skills
Experience with multimodal models such as Qwen-VL, LLaVA, Gemini, GPT vision models, or similar .
Understanding of image preprocessing and document/image understanding.
Experience with OCR, document intelligence, image classification, object detection, or visual question answering is a plus.
Ability to build pipelines combining vision + language + retrieval .