AI Computer Institute
Expert-curated CS & AI curriculum aligned to CBSE standards. A bharath.ai initiative. About Us

The AI Job Market: Career Paths for IIT Graduates

📚 Career Development⏱️ 21 min read🎓 Grade 12
✍️ AI Computer Institute Editorial Team Updated: September 2026 CBSE-aligned · Peer-reviewed · 21 min read
Content curated by subject matter experts with IIT/NIT backgrounds. All chapters are fact-checked against official CBSE/NCERT syllabi.

It is placement season, sixth semester internship season, or the final year — pick whichever applies — and a Computer Science and Engineering student at IIT Bombay is holding four offers. One is "Applied Scientist" at a large e-commerce company. One is "Research Scientist" at a lab that wants a Master's or PhD track. One is "Founding ML Engineer" at a twelve-person seed-stage startup with a large equity number and a small salary. One is "MLOps Engineer" at a cloud infrastructure company. The offer letters all say "AI" somewhere in the title, the CTC figures are all quoted as single numbers, and the student has three days to decide. This is not a hypothetical — it is the actual decision structure that a strong fraction of India's ~17,000 annual IIT graduates across all branches now face, because "AI job" has stopped meaning one thing. It is a family of roles with different daily work, different hiring bars, and — this is the part that trips people up — compensation structures that cannot be compared by reading the top-line number on the offer letter. This chapter builds the taxonomy of what these roles actually are, traces how the IIT hiring pipeline routes students into them, and works through the arithmetic of comparing two real offers correctly, because getting that arithmetic wrong is the single most common and most expensive mistake final-year students make.

Eight roles, two axes

Every job posting with "AI" or "ML" in the title differs along two dimensions that matter more than the title itself: how much of the work is advancing or adapting mathematical methods (research depth) versus how much of the work is building systems those methods run on (systems depth). A useful way to see the whole landscape at once is to plot each role on those two axes, because the axes — not the job title — are what determine the interview you will face and the background you need going in.

AI Career Role Map What the work is actually made of, by role Research track Engineering track Infra / hardware Cross-functional Systems & Production Engineering Depth → Math & Research Depth ↑ Research Scientist Advances SOTA, publishes Entry: MS/PhD + papers Applied Scientist Adapts known ML to product Entry: campus/PPO, no pubs needed AI Accelerator Design GPU / TPU / NPU architecture Entry: EE/dual-degree, VLSI+arch ML Engineer Ships models to production Entry: campus, DSA + ML round MLOps / Platform Eng. Distributed training, GPU infra Entry: campus, strong systems Data Engineer Pipelines feeding ML/analytics Entry: campus, DSA + SQL/systems AI Product Manager Defines roadmap, eval metrics Entry: lateral, 2-3 yrs eng first AI Policy & Governance Compute governance, safety evals Entry: think tank / govt / policy MS Positions are qualitative (typical requirements), not a measured scale. AI PM and Policy also draw on a third axis — business/policy judgment — not shown here.

Research Scientist work is advancing the field itself — new architectures, new training methods, results written up and submitted to venues like NeurIPS, ICML, or ACL. The hiring bar leans on the mathematics an IIT student built up through Grade 12-equivalent depth: linear algebra, optimization, probability, and the ability to read a paper like Vaswani et al.'s 2017 Attention Is All You Need and reproduce its complexity analysis on a whiteboard, not just cite its conclusion. Almost no company hires undergraduates straight into this title without a research track record — a B.Tech project under a faculty member, a publication, or a research internship at a lab such as Microsoft Research India, Google Research India, or Adobe Research counts as that record. This is also the role where the direct-from-IIT path is the exception rather than the rule: a large share of students headed here choose an MS or PhD abroad first, because frontier labs — OpenAI, Google DeepMind, Anthropic, Meta's research groups — hire research scientists overwhelmingly after a doctorate or an exceptional publication history, not through undergraduate campus placement.

Applied Scientist is the role most IIT graduates who want "real ML work" actually land in straight out of campus placement. The job is adapting known techniques — not inventing new ones — to a company's specific data and product: fraud models, recommendation ranking, demand forecasting, personalization. It still draws on real statistics and optimization, but the interview loop tests whether you can correctly apply a method and reason about its failure modes, not whether you can extend the literature.

ML Engineer ships models into production systems — training pipelines, feature stores, serving infrastructure, integration with the rest of the product codebase — with an interview loop that looks like a standard software-engineering loop (data structures and algorithms) plus one ML-specific round. MLOps / ML Platform Engineer goes a layer deeper: building the infrastructure that other ML engineers and scientists run on top of — distributed training clusters, GPU scheduling, experiment tracking, model registries. This is the role where the Grade 11-12 systems and deep-learning foundations compound directly: reasoning about a training job's GPU memory budget means knowing that memory is split across model weights, gradients, optimizer state, and activations, and knowing why techniques like ZeRO (Rajbhandari et al., 2020, ZeRO: Memory Optimizations Toward Training Trillion Parameter Models) or tensor and pipeline parallelism (Shoeybi et al.'s Megatron-LM, 2019; Huang et al.'s GPipe, 2019) exist at all — they exist because a single GPU cannot hold a large model's full training state, and someone has to shard it across a cluster. Data Engineer sits adjacent — less ML-specific math, more schema design, pipeline correctness, and throughput, feeding both product analytics and training sets.

Two roles sit off the pure research/engineering line entirely. AI Product Manager decides what gets built and how success is measured — translating an ambiguous "make search smarter" into a concrete evaluation metric and a prioritized roadmap — and is almost never a fresh-graduate hire; it is a lateral move after two or three years as an engineer or scientist, because you cannot credibly set eval criteria for a model you have never had to debug. AI Accelerator Design works on the chips everything above runs on — GPU, TPU, and NPU architecture at companies like Nvidia, Qualcomm, or AMD, or within India's own semiconductor push. It draws specifically from Electrical Engineering and dual-degree students, and it is a direct extension of exactly the kind of reasoning this course's transformer-internals unit covers: knowing that transformer inference at small batch sizes is memory-bandwidth-bound rather than compute-bound is precisely the fact an accelerator architect needs to design around. AI Policy & Governance — at government bodies working on India's compute and AI-safety policy, or within a large lab's policy team — needs far less code but real fluency in compute scaling, capability evaluation, and the economics of export controls on advanced accelerators; it is a small but growing track, usually reached via a policy-focused Master's rather than direct placement.

How the IIT pipeline actually routes you into these roles

The placement process itself shapes which of these roles a given student is even likely to see. Institute placement cells run a slotted process — a pre-placement talk, an online assessment (data structures and algorithms, sometimes with a statistics or ML section for AI-titled roles), then technical interview rounds — with companies grouped into slots by role seniority and package, so the most competitive AI roles cluster in the earliest slots and a student who accepts an offer there typically exits the pool for everyone after. Running in parallel to full-time placement is the summer internship track after the sixth semester: a strong two-to-three-month internship performance frequently converts into a pre-placement offer (PPO), which is how a large share of AI-track full-time offers at major recruiters actually originate — students who interned well simply never enter the main placement round. Research-track hiring runs on a different clock entirely. It rewards an undergraduate research record built over multiple semesters — a thesis under a faculty member in one of the institute's AI-focused groups, a research internship, ideally a paper — and because that record takes years to build and the strongest research roles want a doctorate, many research-inclined students opt out of the direct-placement pipeline altogether and apply to MS/PhD programs instead, landing at a frontier lab only after that.

The misconception: comparing offers by CTC

The single most expensive mistake in this whole process is comparing offers by their headline Cost-to-Company number. CTC bundles unvested, sometimes illiquid equity into an annualized figure as if it were already cash in hand — but real equity grants vest over years (often with a cliff, meaning zero if you leave the company before the cliff date), and a startup's equity is not a guaranteed payout at all; it is a claim on a company that might be worth nothing in a few years. The correct comparison is a discounted, probability-weighted cash-flow analysis — an NPV calculation — not the number printed at the top of the offer letter.

Worked example: two offers, one correct comparison

An IIT graduate has two offers, both illustrative:

Offer A — Applied Scientist, established company. Base ₹28 lakh/year, a ₹5 lakh joining bonus (Year 1 only), and an RSU grant worth ₹40 lakh at offer time, vesting back-loaded over four years at 10% / 20% / 30% / 40% — a common structure companies use specifically to improve later-year retention, rather than an even 25% split.

Offer B — Founding ML Engineer, seed-stage startup. Base ₹18 lakh/year, no bonus, and equity of 1.2% of a company currently valued at ₹80 crore post-money — a nominal ₹96 lakh (1.2% of ₹80,00,00,000 = ₹96,00,000). That number is meaningless on its own: it only pays out if the company survives to an exit. Assume, for this analysis, a 15% probability the startup reaches an exit where the equity retains its full nominal value, materializing at the end of Year 4; otherwise it is worth zero. Its expected value is therefore 0.15 × ₹96 lakh = ₹14.4 lakh, not ₹96 lakh.

Holding raises flat for both offers (a simplifying assumption, stated explicitly so it doesn't get mistaken for a real prediction), the year-by-year cash flows in ₹ lakh are:

Offer A: Year 1 = 28 + 5 + 4 (10% of 40) = 37. Year 2 = 28 + 8 (20% of 40) = 36. Year 3 = 28 + 12 = 40. Year 4 = 28 + 16 = 44.
Offer B: Year 1–3 = 18 each (base only). Year 4 = 18 + 14.4 (expected equity) = 32.4.

Discounting both at 8% — a rate reflecting the time value of a fairly liquid, low-idiosyncratic-risk stream, since the startup's default risk is already handled by probability-weighting the equity rather than by inflating the discount rate a second time — gives the net present value:

def npv(cash_flows, rate):
    return sum(cf / (1 + rate) ** t for t, cf in enumerate(cash_flows, start=1))

offer_a = [37, 36, 40, 44]     # base + bonus + vested RSU, ₹ lakh, Year 1-4
offer_b = [18, 18, 18, 32.4]   # base only, then base + expected equity, Year 1-4

print(round(npv(offer_a, 0.08), 1))  # 129.2
print(round(npv(offer_b, 0.08), 1))  # 70.2

Tracing it by hand confirms the code: Offer A's four discounted terms are 37/1.08 ≈ 34.26, 36/1.1664 ≈ 30.86, 40/1.259712 ≈ 31.75, 44/1.36048896 ≈ 32.34, summing to ≈ ₹129.2 lakh. Offer B's terms are 18/1.08 ≈ 16.67, 18/1.1664 ≈ 15.43, 18/1.259712 ≈ 14.29, 32.4/1.36048896 ≈ 23.82, summing to ≈ ₹70.2 lakh. Despite Offer B's headline equity number (₹96 lakh) looking larger than Offer A's entire RSU grant (₹40 lakh), Offer A's risk-adjusted, time-adjusted value is nearly double Offer B's. That gap is the whole lesson: a large nominal equity number at a low-survival-probability company is worth far less than it appears, and the correct comparison discounts for both risk (via the survival probability) and time (via the discount rate) — never just one.

Why the frontier-lab hiring bar goes past this

For research-track roles specifically, the compensation analysis above is table stakes; the technical bar sits on production-systems literacy this course has already built toward. An interview for a research or applied-science role at a lab training large models expects you to reason about why the Chinchilla scaling result (Hoffmann et al., 2022, DeepMind) implies a specific compute-optimal ratio between model size and training tokens, not just that "bigger models need more data." An MLOps interview expects you to reason about why a 70-billion-parameter model's optimizer state alone (in mixed-precision Adam, roughly 12 bytes per parameter across master weights, momentum, and variance) will not fit on a single 80GB GPU, and why that forces the sharding strategies ZeRO and Megatron-LM were built to solve. These are not decorative facts — they are the actual content of the interview loop at that end of the market, and they are exactly the kind of question a headline CTC number gives you no preparation for.

Active recall

Attempt each question before reading its answer.

1. A student has a research internship, no publication, and wants a Research Scientist offer directly from IIT campus placement without further study. How realistic is this, and what would most plausibly get them there instead?

2. Suppose the startup in the worked example raises a strong Series A, and its probability of a successful exit is revised from 15% to 35%. Recompute Offer B's Year 4 cash flow and its total NPV at 8%, and state whether the ranking against Offer A flips.

3. Using the original 15% exit probability, recompute both offers' NPV at a 15% discount rate instead of 8%. Does the ranking change? Which offer's NPV falls by more in absolute ₹ lakh, and why does that make structural sense given each offer's cash-flow shape?

4. Why does comparing two offers by their stated CTC systematically favor offers with large illiquid or unvested components, even when those offers are worse in expected-value terms?

5. A student values learning speed and founder-level ownership more than expected monetary value, and is aware Offer B's NPV is lower. Under what conditions is choosing Offer B still a rational decision?

6. An MLOps Engineer candidate is asked why a 70B-parameter model's training memory footprint cannot simply be reduced by using a bigger single GPU. What is the correct answer?

Answers.

1. Unlikely without further study. Direct-from-IIT Research Scientist offers are the exception; the role's hiring bar is built on a multi-year research record (thesis, publications, sustained work with a faculty group), which an internship without a publication does not yet establish. The more plausible path is an MS or PhD — often funded partly by that same internship's letters of recommendation — followed by a research-scientist offer post-degree, or an Applied Scientist offer now with a lateral move into research later once a stronger track record exists.

2. New expected equity value = 0.35 × ₹96 lakh = ₹33.6 lakh, so Year 4 cash flow becomes 18 + 33.6 = ₹51.6 lakh. Its discounted term is 51.6/1.36048896 ≈ 37.93. Only the Year 4 term changes; Years 1–3 stay at 16.67 + 15.43 + 14.29 ≈ 46.39. New total NPV ≈ 46.39 + 37.93 = ₹84.3 lakh. Offer A's NPV (₹129.2 lakh) is unchanged since nothing about Offer A was touched. The ranking does not flip — Offer A still dominates, though the gap narrows from about ₹59 lakh to about ₹45 lakh.

3. At 15%, Offer A's NPV becomes ≈ 32.17 + 27.22 + 26.30 + 25.16 ≈ ₹110.9 lakh, and Offer B's (at the original 15% exit probability) becomes ≈ 15.65 + 13.61 + 11.84 + 18.52 ≈ ₹59.6 lakh. The ranking is unchanged — Offer A still wins. Offer A's NPV falls by more in absolute terms (₹129.2 → ₹110.9, a drop of about ₹18.3 lakh) than Offer B's (₹70.2 → ₹59.6, a drop of about ₹10.6 lakh), because Offer A has larger absolute cash flows in every year being discounted, so the same percentage discount removes a larger rupee amount from it.

4. CTC is typically computed by annualizing the full grant value at the moment of offer, as if 100% had already vested and as if illiquid equity were guaranteed cash. That inflates any offer with a large equity or bonus component relative to one weighted toward guaranteed base salary, exactly backwards from what expected-value analysis shows once vesting schedules, cliffs, and survival probability are correctly accounted for.

5. It's rational when the decision-maker correctly values something NPV doesn't price: faster ownership of ambiguous, high-stakes problems (a genuine career-capital return not captured in Year 1–4 cash flow), a personal risk tolerance that makes a low-probability, high-upside outcome worth more subjectively than its expected value, or diversification — if this is one bet among several career moves rather than a one-shot decision, taking the higher-variance option once can be rational even at a lower expected value.

6. Training memory isn't just the model's weights — it also holds gradients and, for an optimizer like Adam, per-parameter momentum and variance terms, plus activations kept for backpropagation. For 70 billion parameters, that state alone runs to roughly 12 bytes per parameter even before activations, which exceeds what any single GPU (typically 80GB) can hold. The fix isn't a bigger single chip — it's sharding that state across many GPUs, which is exactly what techniques like ZeRO and tensor/pipeline parallelism (Megatron-LM, GPipe) are built to do.

Think About It

Think about this: How would you explain the ai job market: career paths for iit graduates to a friend who has never seen a computer? What real-world analogy would you use? Imagine you had to build a system using these concepts — what would be your first step? Try this: before moving on, write down three things you learned and one question you still have.

Practice Exercises

Now it is time to practice! Complete these challenges to solidify your understanding:

  • Exercise 1: Write a short program that demonstrates the core concept from this chapter. Test it with at least 3 different inputs.
  • Exercise 2: Find a real-world example where the ai job market: career paths for iit graduates is used in an Indian company (like TCS, Infosys, Flipkart, or ISRO). Write a paragraph explaining the connection.
  • Exercise 3: Create a mind-map connecting the ai job market: career paths for iit graduates to at least 3 other topics you have studied.

Key Takeaways — Summary and Recap

Let us recap what we covered: the core ideas behind the ai job market: career paths for iit graduates, how they connect to real-world applications, and why they matter for your journey in computer science. Remember these key points as you move forward. For competitive exam preparation (CBSE, JEE, BITSAT), focus on understanding the WHY behind each concept, not just the WHAT.

← AI Startups: Building an AI Company in IndiaOpen Source AI: Hugging Face, LangChain, and the Ecosystem →

Found this useful? Share it!

📱 WhatsApp 🐦 Twitter 💼 LinkedIn