Skip to main content
E

ekacare

Senior Data Scientist ( Kernel Optimisation & Inference Engineer)

Bengaluru, India2-5 yrsPosted on Aug 13, 2026

Skills

GoRubyNode.jsCommunication

Job description

Kernel Optimisation & Inference Engineer

Bengaluru · Full-time · 2–5 yrs

Somewhere between the model and the silicon, 10–20% of a training budget goes missing. Your job is to go get it back, and then make the same model fast enough to run in a clinic, and small enough to run on a phone.

About EkaCare and the mission

EkaCare is India's connected healthcare platform: an EMR that doctors run their practices on, a personal health record used by millions of Indians, and one of the deepest integrations with India's ABDM digital-health rails. Our Parrotlet family of medical models already serves Indian doctors in production, and we open-source our work where it counts.

The role

You'll work with our performance lead on making everything fast: training-side fused kernels and MFU on the 30B MoE, inference-side latency and throughput, and the quantised 2B/4B on-device tier. Hardware-up: profiler first, roofline reasoning always, custom kernels when the math says so.

What you'll do

Profile training and inference workloads and hunt utilisation gaps across kernels, memory and comms.

Write and tune CUDA kernels where existing ops leave real performance on the table, and know when they don't.

Optimise MoE-specific paths: grouped GEMMs, all-to-all communication, expert load imbalance.

Build the fast inference path: vLLM-class serving, continuous batching, prompt/prefix caching for clinical-context workloads, speculative decoding.

Own quantisation for the 2B/4B variants (AWQ/GPTQ-class, fp8) — with eval-parity verification, not just perplexity.

Make on-device inference real for the hardware Indian clinics actually have.

What we look for

2–5 years in GPU performance work; you've profiled real workloads and shipped optimisations with before/after numbers you can defend.

Working fluency in CUDA, and memory-hierarchy reasoning (coalescing, occupancy, SRAM tiling; you can explain *why* FlashAttention is fast).

Hands-on with a modern serving stack (vLLM, TensorRT-LLM, SGLang or similar) beyond just running it.

Measurement discipline: you profile before optimising and verify correctness after.

Bonus

fp8 experience on H100/H200-class hardware; torch.compile/inductor internals.

Quantisation research or on-device/mobile inference experience.

Open-source kernels or serving contributions.

Why this is a rare gig

Open source, with your name on it : weights and technical reports ship publicly.

India-scale mission : models for a billion people in their own languages.

Compute that’s rare to fine : dedicated multi-node H200 training under a national grant.

Small senior team: you work with the people who own the recipe.

A live deployment path : Government institutes, EkaCare's doctors and patients use what you ship.

Full-Time Employee Benefits

Medical Insurance & Accidental Insurance

Maternity & Paternity Benefits

PF, Gratuity, & Leave Encashment

Salary Advance Policy

Apply on ekacare