ekacare
Senior Data Scientist ( Kernel Optimisation & Inference Engineer)
Skills
Job description
Kernel Optimisation & Inference Engineer
Bengaluru · Full-time · 2–5 yrs
Somewhere between the model and the silicon, 10–20% of a training budget goes missing. Your job is to go get it back, and then make the same model fast enough to run in a clinic, and small enough to run on a phone.
About EkaCare and the mission
EkaCare is India's connected healthcare platform: an EMR that doctors run their practices on, a personal health record used by millions of Indians, and one of the deepest integrations with India's ABDM digital-health rails. Our Parrotlet family of medical models already serves Indian doctors in production, and we open-source our work where it counts.
The role
You'll work with our performance lead on making everything fast: training-side fused kernels and MFU on the 30B MoE, inference-side latency and throughput, and the quantised 2B/4B on-device tier. Hardware-up: profiler first, roofline reasoning always, custom kernels when the math says so.
What you'll do
Profile training and inference workloads and hunt utilisation gaps across kernels, memory and comms.
Write and tune CUDA kernels where existing ops leave real performance on the table, and know when they don't.
Optimise MoE-specific paths: grouped GEMMs, all-to-all communication, expert load imbalance.
Build the fast inference path: vLLM-class serving, continuous batching, prompt/prefix caching for clinical-context workloads, speculative decoding.
Own quantisation for the 2B/4B variants (AWQ/GPTQ-class, fp8) — with eval-parity verification, not just perplexity.
Make on-device inference real for the hardware Indian clinics actually have.
What we look for
2–5 years in GPU performance work; you've profiled real workloads and shipped optimisations with before/after numbers you can defend.
Working fluency in CUDA, and memory-hierarchy reasoning (coalescing, occupancy, SRAM tiling; you can explain *why* FlashAttention is fast).
Hands-on with a modern serving stack (vLLM, TensorRT-LLM, SGLang or similar) beyond just running it.
Measurement discipline: you profile before optimising and verify correctness after.
Bonus
fp8 experience on H100/H200-class hardware; torch.compile/inductor internals.
Quantisation research or on-device/mobile inference experience.
Open-source kernels or serving contributions.
Why this is a rare gig
Open source, with your name on it : weights and technical reports ship publicly.
India-scale mission : models for a billion people in their own languages.
Compute that’s rare to fine : dedicated multi-node H200 training under a national grant.
Small senior team: you work with the people who own the recipe.
A live deployment path : Government institutes, EkaCare's doctors and patients use what you ship.
Full-Time Employee Benefits
Medical Insurance & Accidental Insurance
Maternity & Paternity Benefits
PF, Gratuity, & Leave Encashment
Salary Advance Policy