Certification in LLM Inference Engineering & AI Infrastructure

Build the systems that keep LLMs fast, reliable and in production.

Offline, taught in person Masai HQ, Bengaluru

Next cohort starts November 21st

20
Weekends
240
Learning hours
4
Projects
2+ yrs
Experience
The shift

How AI is reshaping the industry

Every industry is being rebuilt around models. The engineers who can run them in production are the ones doing the rebuilding.

Adaptation
Fine-tuning
Attention
Tokenization
Evaluation
LLM-as-judge
Benchmarking
Inference efficiency
KV caching
Continuous batching
Quantization
Speculative decoding
Tensor parallelism
GPU profiling
Serving
Autoscaling
Observability
Reliable operations
Capacity and cost
Paper implementation

Scroll to explore

Phase-by-phase curriculum

What you'll learn

5 phases across 20 weekends — scroll to move through them.

Phase A 01 / 05
Weeks 1–4

Model and Measurement Foundations

Understand how a model executes, then establish the quality and performance baselines every later optimization is measured against.

  1. Week 1: LLM execution under the hood Understand how an LLM actually works inside.
  2. Week 2: Autoregressive inference mechanics See where every millisecond of a response goes.
  3. Week 3: Evaluation and quality baselines Decide what "good output" means, and measure it.
  4. Week 4: Performance benchmarking and profiling Put hard numbers on speed, cost and throughput.
Phase B 02 / 05
Weeks 5–6

Hardware Prerequisite

Reason from GPU compute, memory and interconnect constraints before touching engine internals.

  1. Week 5: GPU compute architecture for inference Learn what the GPU is really doing.
  2. Week 6: GPU memory, interconnects and topology Know where memory goes and when it runs out.
Phase C 03 / 05
Weeks 7–14

Inference Engines and Optimization

One continuous block from engine internals to measured single-node optimization — the core of the program.

  1. Week 7: vLLM deep dive I: serving and scheduling Swap the naive server for a real one.
  2. Week 8: vLLM deep dive II: PagedAttention and KV memory Make the KV cache stop wasting memory.
  3. Week 9: SGLang deep dive I: runtime and RadixAttention Reuse what the model has already computed.
  4. Week 10: SGLang deep dive II: workload-aware optimization Tune the engine to the traffic you actually get.
  5. Week 11: TensorRT-LLM, deployment artifacts and Triton Ship a compiled, optimized artifact.
  6. Week 12: Quantization and precision engineering Shrink the model without breaking it.
  7. Week 13: Speculative decoding in depth Generate several tokens for the price of one.
  8. Week 14: Prefix, index and tiered KV-cache systems Cache across requests, GPUs and storage.
Phase D 04 / 05
Weeks 15–18

Distributed Inference and Platform

Scale the runtime past one GPU, then build the serving platform that deploys, routes, observes and recovers it.

  1. Week 15: Multi-GPU, MoE and distributed inference Scale past a single GPU.
  2. Week 16: Prefill-decode disaggregation and routing Split prefill from decode, and route by cache.
  3. Week 17: Platform engineering and model lifecycle Get it from a registry to a real endpoint.
  4. Week 18: LLMOps, traffic management and reliability Keep it up when traffic spikes and things fail.
Phase E 05 / 05
Weeks 19–20

Bounded Extensions and Capstone

Add the adaptation and application layers the job needs, then defend the whole system in front of a panel.

  1. Week 19: Bounded adaptation and application workloads Fine-tune, serve adapters, and wire up retrieval.
  2. Week 20: Capstone validation and technical defence Prove the whole thing works, and defend it.
Who teaches this

No trainers. Only engineers who ship this.

They teach each week and review your code. They have already solved the problems you are about to hit.

  • Tarun Tiwari

    Tarun Tiwari

    Head of Product, Nava

    Built the Alexa AI platform behind 100M+ devices. Now runs GPU cloud and training clusters.

    GPU cloud & infrastructure Shipped systems at Alumni of
  • Ashish Prasad

    Ashish Prasad

    Engineering Manager (AI/ML), Microsoft Turing

    Helped put GPT-5 inside Microsoft 365 Copilot, working with a team of 30 across engineering and research.

    LLM systems & retrieval Shipped systems at Alumni of
  • Shubhendu Shishir

    Shubhendu Shishir

    Head of Engineering, Simplismart

    Runs the engineering behind one of the fastest inference engines in production.

    Inference & serving Shipped systems at Alumni of
  • Naman Vats

    Naman Vats

    Senior Software Engineer, Sentient Labs

    Published research on why the same AI model can cost 40× more to run depending on how you build around it.

    Agents & evaluation Shipped systems at Alumni of
  • Mayank Sharma

    Mayank Sharma

    AI Lead, Opkey

    Trains models from scratch, and has led ML teams building for images, language and on-device AI.

    Model training & fine-tuning Shipped systems at Alumni of
Program at a glance

LLM Systems is the field; Infra is the specialization.

You learn the model and the system that runs it, end to end.

Duration

20 weekends

Across at least five months.

Experience

2+ years

Relevant full-time work; 3+ preferred.

Guided hours

240 hours

Plus 80–120 independent; 320–360 total.

Weekend format

6 + 6 hrs Sat / Sun

Live weekend sessions, plus about 30 minutes a day of weekday reading.

Who it is for

ML and applied ML engineers who already work in PyTorch, and backend or platform engineers who meet the same bar.

Eligibility

Can you join? Here is what we look for.

Your job title does not matter. If most of these sound like you, apply.

  1. 2+ years, full-time Relevant professional work. Internships do not count.
  2. ML or applied ML background Or backend/platform, if you meet the same ML bar.
  3. Professional Python Modular code, testing, debugging, Git.
  4. Hands-on PyTorch Run and modify a model, inspect tensors, use autograd.
  5. Transformers at working level Attention, embeddings, tokenization, how generation works.
  6. Working systems knowledge Linux, HTTP, concurrency, Docker.
Time commitment

How the weekend is structured

Saturday6h

2.5 hours instruction, 3 hours guided lab, 0.5 hour checkpoint.

Sunday6h

2.5 hours workshop, 3 hours build/benchmark task, 0.5 hour review.

Weekdays30 min/day

Pre-read and experiment notes between sessions.

Program Promise
You will understand an open LLM, make it faster and cheaper, and run it as a reliable production service.
Program fee

Program fee and what it covers.

Pay the application fee to secure your interview slot. The rest is due only once you have an offer and have accepted your seat.

What the fee covers

  • 240 guided hours across 20 weekends
  • Live weekend sessions with practitioner instructors
  • GPU compute for every lab, benchmark and capstone run
  • Capstone mentorship through to submission
Application fee 100% refundable if not shortlisted for the interview ₹999
Total program fee

1,75,000+ 18% GST

  1. Registration fee non-refundable ₹5,000
  2. Program fee refundable as per refund policy ₹1,70,000
  3. GST ₹31,500

Seats are confirmed in order of acceptance.

Pay Your Program Fee, Your Way

Best Value

Upfront Payment

Pay once, save more.

₹2,06,500

Pay In Parts

Flexible plans, no financial stress.

Instalment 1 ₹50,000
Instalment 2 ₹1,56,500

EMI via NBFC Partners

Spread the cost over 6 months.

₹37,613 / month × 6

Inclusive of GST

Program fee is refundable as per our Refund & Withdrawal Policy

Twenty weekends. Four Projects. One production-grade skill set.

Apply Now
Questions

The things people ask before applying

Do not see your question? Ask us before you apply.

What does the fee include?

All 240 guided hours, live weekend sessions, project reviews, capstone mentorship — and the GPU compute you use throughout. Every lab, benchmark and capstone run happens on machines we provide, so there is no separate cloud bill. You just need your own laptop.

When do I pay?

In three stages. ₹999 when you apply, which secures your interview slot. If you are selected, a registration fee of ₹5,000 confirms your seat. The balance is due before the first weekend, and you can pay it upfront, in two instalments, or as 6 monthly EMIs through our NBFC partners.

Do I need a machine learning background?

Yes. This one is ML-first: you should already write production Python, work comfortably in PyTorch, and understand how transformers generate text. Backend and platform engineers are welcome, but they meet the same ML bar rather than substituting systems experience for it.

How much time does it take?

Six hours on Saturday, six on Sunday, plus about 30 minutes a day during the week. About 320–360 hours in total.

Can I keep my full-time job?

Yes. That is what the weekend format is for — we expect everyone to be working, and the weekday work is built to fit around a job.

What if I miss a weekend?

Sessions are recorded, so missing one is fine. Miss several in a row and your project deadlines slip, so we would work out a plan with you.

What do I finish with?

You will understand an open LLM, make it faster and cheaper, and run it as a reliable production service.

Is there a placement guarantee?

No. This is for engineers who already have jobs. What you get is proof you can run LLM systems in production — the four projects.