vLLM x NVIDIA Dynamo Meetup
- WHO IT IS FOR: This event is designed for engineers, researchers, and builders focused on LLM serving and inference optimization.
- WHAT IT IS ABOUT: The meetup explores efficient LLM serving at scale, focusing on the latest developments from vLLM and NVIDIA Dynamo.
- FORMAT: The evening features a series of technical talks followed by a networking social with food and drinks.
Join the vLLM community and the NVIDIA Dynamo team in San Francisco for an exclusive evening of technical insights and networking. This event dives deep into the complexities of serving Large Language Models efficiently at scale, covering critical topics such as inference optimization, distributed serving, and the real-world challenges of production deployments.
The evening kicks off with a series of focused tech talks from the vLLM and Dynamo teams, providing a glimpse into the cutting-edge work pushing the boundaries of AI inference. Whether you are an expert in inference engine internals or a developer new to LLM serving, this is a prime opportunity to learn from the architects of these systems.
Following the presentations, guests are invited to connect with fellow researchers and builders from across the inference community. Enjoy food and drinks while engaging in high-level discussions about the future of AI infrastructure. Space is limited and registration is subject to approval, so be sure to reserve your spot early.