Abi Aryan

Abi Aryan

1.6K Subscribers

Computer Scientist and ML Engineer

Maven's Terms and Privacy Policy.

Abi Aryan Inference Infra Queen (GPU Infra & LLM Serving)

Abi Aryan is the ex-founder of Abide AI and a machine learning engineer with over a decade of experience building production-level ML systems.

A mathematician by training, she previously was a visiting research scholar at UCLA, supervised by the 2012 Turing Award winner, Dr. Judea Pearl, where she focused on developing intelligent agents.

Abi has authored research papers in AutoML, multi-agent systems, and large language models, and actively reviews for leading research conferences and workshops, including NeurIPS, ACL, EMNLP, and AABI

She is currently teaching and advancing doctoral research at the intersection of several key areas:

• Distributed systems

• GPU engineering for large-scale AI systems

• SLO-Aware Inference Optimization

• Low-cost chip design

• AI Policy

Book Author (LLMOps, GPU Engineering for AI Systems)
O'Reilly Media
@Packtpub

Alumni reviews

Great content and assignments!

Piotr

Pikachu
AI Engineer · praktika.ai
This course provides a practical walkthrough of LLM inference, deployment, and scaling. I especially liked the focus on production systems and engineering trade-offs, along with useful guest lectures that added industry perspective. By the end, the course builds solid intuition for system design and deploying LLMs at scale. Highly recommend.

Vikram

Pikachu
Cloud Engineer · Manulife
Fantastic course for deep-diving into LLM systems. Each week features a great mix of lectures and code demos, supported by detailed documentation. The final project is a highlight, covering the full inference lifecycle: load balancing, agentic systems, and observability with metrics/traces. Abi is a dedicated instructor who truly listens to student feedback. The 1:1 career guidance and system design sessions make this an incredible value for infrastructure-focused engineers.
Reviewer profile image

Vivian

Pikachu
Software Engineer · Google
Abi Aryen is an exceptional instructor and mentor. From the very first class, she demonstrated genuine enthusiasm and generosity with her time, ensuring every student felt supported. When many of us asked for additional sessions, Abi didn’t hesitate—she was genuinely excited to teach more and help us dive deeper into the material. The course content itself was outstanding. It’s not something you can find neatly packaged anywhere else right now. The material offers deep insight into how modern inference systems work, how to build effective mental models around performance bottlenecks, and how to reason systematically about optimization challenges in real-world environments. One of the most valuable aspects of the course was the hands-on project. Over seven weeks, we each designed and implemented our own inference gateway—a fantastic “learn by doing” approach that really cemented the concepts covered in class. By the end of the course, every student had a fully functional gateway that can operate across local and cloud inference providers, integrating seamlessly with virtually any inference engine. It even includes robust metrics tracking across the entire inference infrastructure, making the project practical for both homelab and production-level environments. Finally, Abi’s dedication to her students goes far beyond expectations. She routinely held extra office hours and provided personal guidance throughout the week. It’s rare to meet someone as generous with their time and expertise as Abi. This course is a must for anyone serious about mastering AI systems design and inference engineering.

Shawn

Pikachu
Researcher · Palo Alto Networks
Thank you to Abi Aryan for designing and teaching this course. Most inference content stops at "run vllm serve and benchmark it." This one went deeper: routing layers, experiment controls, observability as a first-class concern, agentic workloads as a distinct traffic class. The InferenceOps framing - owning the full path from the first request byte to the served token - is something I'll carry into production work from here. The finding that changed how I think about benchmarking came directly from the assignment structure. Running gateway-direct and agent end-to-end as separate load-test tiers revealed a 2.6× p95 divergence between baseline and chunked prefill that a single-call benchmark would have buried entirely. That's not a measurement trick it's a mental model shift about what "latency" means when the caller is an agent making sequential LLM calls, not a human waiting on response. To the cohort: the discussions around speculative decoding tradeoffs, KV cache behaviour under concurrency, and cost proxies were the kind of pressure-testing you can't manufacture alone. Good people to build alongside. Right course, right time.

Karthik

Pikachu
AI Engineer · CCC