Best Free Resources (better than paid ones)
Mathematics (Must):
For intuition behind LLMs, Embeddings use Linear Algebra and Gradient Descent uses Calculus.
Even if you have done Math in collage, retake these for visual understanding otherwise you will not understands LLMs properly.
Linear algebra
https://www.youtube.com/playlist?list=PLZHQObOWTQDPD3MizzM2xVFitgF8hE_ab
Calculus
https://www.youtube.com/playlist?list=PLZHQObOWTQDMsr9K-rj53DwVRMYO3t5Yr
Deep Learning (2 alternatives)
Overview Series: MIT
https://www.youtube.com/playlist?list=PLWxrjrFiQajE
(Optional)
Only for advance learner
Mitesh Khapra (IIT Madras) 🧠
https://www.youtube.com/playlist?list=PLZ2ps__7DhBZVxMrSkTIcG6zZBDKUXCnM
LLM (all track must except theory for advance learner)
Visual (3Blue1Brown)
https://www.youtube.com/playlist?list=PLZHQObOWTQDNU6R1_67000Dx_ZCJB-3pi
Theoretical (optional) (Mitesh Khapra 🧠) for advance learner but GOLD resource
https://www.youtube.com/playlist?list=PLZ2ps__7DhBbaMNZoyW2Hizl8DG6ikkjo
Practical (Andrej Karpathy 🧠)
https://www.youtube.com/playlist?list=PLAqhIrjkxbuWI23v9cThsA9GvCAUhRvKZ
Pytorch (optional)
https://www.youtube.com/watch?v=V_xro1bcAuA&t=78935s
After hands-on above, do Full Course here, with advanced projects:
https://learnpytorch.io/
Learn about-
Tokenization, context windows, few-shot prompting, structured outputs.
Tokenization- track token usage and latency
Context window - how much info model can consider, what happen if they exceed it
Temperature and other parameter - how affect randomness and creativity
Structured outputs - clean json, response format, output validation, handle when model ignore instrution
Streaming generation, error handling when api fail, rate limits, retry logic, cost management
Tool use (function calling), tool chaining. Build multi-step applications with real actions.
Tools Schema, Augment Validation, Failure handling, Guardrails
Zero-shot, few-shot prompting
• Chain-of-thought prompting
• KV cache and how it works
• Prompt caching: what it is and when it cuts costs
Course
Transformer in Practice (VP @ AMD 🧠)
https://www.deeplearning.ai/courses/transformers-in-practice
Hugging Face (must)
https://www.youtube.com/watch?v=XNDY7jtihok&t=7910s
After hands-on above, do Full Course here, with advanced projects
https://www.learnhuggingface.com/
LangChain
https://www.deeplearning.ai/courses/langchain
Advance
https://www.deeplearning.ai/courses/functions-tools-agents-langchain
Project-
Domain specific assistance, structured data extractor.
Research assistance that search's web, read files, call an API.
AI Engineering Book by Chip Huyen (Bible for AI Engineering)
Only book you need to read for AI Engineering
Link: https://drive.google.com/file/d/1jfn7vc8yVW7S-fToWVotjeRYuoildK29/view?usp=sharing
Learn about-
Chunking, embeddings, vector databases, retrieval strategies.
What retrieval-augmented generation actually solves
Chunking strategies, avoid retrival modes, configure mode and embedding models.
Context assembly problem
Retrieval strategies - similarity, hybrid, reranking, query transformation
Learn these Framework
Pgvector - index and distance metrics choises, hybrid retrieval pattens, Chunking failure modes
weaviate - Hybrid retrival, schema design, ingestion pattens, index configuration pattens, multi-tetants pattens
Course (weaviate included)
https://www.deeplearning.ai/courses/retrieval-augmented-generation
pgvector
https://www.youtube.com/watch?v=hAdEuDBN57g
Prompt Compression and Query Optimization
https://www.deeplearning.ai/courses/prompt-compression-and-query-optimization
AI Agent
Learn about-
Agent Loops & Protocols
Core agent loop, planning, memory, error handling.
Agent Swarms, Self improving Agents, RL for Agents.
State Machine, Planning vs execution, memory management, obserbility.
Chain of thought, Tree of thought.
Episodic vs semantic memory. Red teaming. Cost Saving.
Tool use and function calling
ReAct pattern: Reasoning + action (Thought → Action → Observation loop)
Multi-step planning and agent loops
Where agents break in production.
Build your own Claude Code Agent with Python: Learn about LLM APIs, tool calling, agent loops and more
https://app.codecrafters.io/courses/claude-code/overview
LangGraph (must)
https://www.deeplearning.ai/courses/long-term-agentic-memory-with-langgraph
Pydantic for LLM Workflows
https://www.deeplearning.ai/courses/pydantic-for-llm-workflows
Building Your Own Database Agent
https://www.deeplearning.ai/courses/building-your-own-database-agent
Building and Evaluating Data Agents
https://www.deeplearning.ai/courses/building-and-evaluating-data-agents
Project-
Multi turn conversation app (conversation history, context management, iteration workflow)
Multi AI Agent (Optional)
https://www.deeplearning.ai/courses/practical-multi-ai-agents-and-advanced-use-cases-with-crewai
https://www.deeplearning.ai/courses/building-ai-voice-agents-for-production
Google’s ADK Live Voice Agents (Optional)
https://www.deeplearning.ai/courses/building-live-voice-agents-with-googles-adk
MCP
Least privilege tool exposure, auditing
Start here
https://anthropic.skilljar.com/introduction-to-model-context-protocol
Advance
https://anthropic.skilljar.com/model-context-protocol-advanced-topics
Claude Code in CLI
Start here
https://anthropic.skilljar.com/claude-code-in-action
Claude Sub-agents
https://anthropic.skilljar.com/introduction-to-subagents
Agent Skills
https://anthropic.skilljar.com/introduction-to-agent-skills
Learn about-
Metrics like Faithfulness, relevance, groundedness. Find which is relevent to your use case and why.
Build dataset from real dataset. Run offline vs online evals workflows. Use diff evaluator types - rule based/LLM as a judge/human review.
Build evaluation sets, Run regression test on prompt, track latency, cost and error rates. Set up alerts.
Regression testing for retrival and prompt changes.
Observability, Tracing and monitoring. Trace level debugging, comparing releases on latency, cost and quality. Linking evaluation regression back to concrete traces.
Learn these Framework
Langsmith - Native observability and evaluation.
ragas - LLM-as-a-Judge
Learn about-
FastAPI, Docker, cloud deployment, logging, observability, cost tracking, caching.
Container, Async execution, background jobs, latency monitoring.
Cost tracking - per request, day, user. Caching aggressively, batch request, use smaller models when possible. Health checks, basic alerting, optimize cost.
High performance Inference. For embedding, rerankers, cv models.
Nvidia Triton server, Bento ML. Dynamic patching, concurancy tuning, model repository, Ensembles for pipeline serving.
Google LLMOps
https://www.deeplearning.ai/courses/llmops
https://www.deeplearning.ai/courses/fast-and-efficient-llm-inference-with-vllm
Safe and reliable AI via guardrails
https://www.deeplearning.ai/courses/safe-and-reliable-ai-via-guardrails
Quantization Fundamentals with Hugging Face (Optional)
https://www.deeplearning.ai/courses/quantization-fundamentals
Quantization in Depth (Optional)
https://www.deeplearning.ai/courses/quantization-in-depth
Fine Tuning (Optional)
Learn about-
PEFT & Quantization Strategies
When to fine-tune vs prompt engineer.
PEFT: parameter-efficient fine-tuning
LoRA and QLoRA (both are PEFT methods, QLoRA adds 4-bit quantization on top)
Instruction tuning vs continued pre-training
How much throughput
Tail latency work.
KV cache constrints.
Memory behaviour.
Quantization tradeoff.
Course
Google RLHF (Optional)
https://www.deeplearning.ai/courses/reinforcement-learning-from-human-feedback
Interview: Include architecture diagram, discuss tradeoff, talk about what broke, and how you fixed it.
Go deep into AI specialisation.
More latest resources will be added as required in future.