From Vision-Language Models to Physical AI: Embedded Intelligence Enters a New Phase

Heading into this year’s Embedded Vision Summit, the most important story in embedded AI is not just that systems are getting smarter, but that they are becoming more practical. Embedded AI is moving beyond recognition toward systems that can better…

Heading into this year’s Embedded Vision Summit, the most important story in embedded AI is not just that systems are getting smarter, but that they are becoming more practical. Embedded AI is moving beyond recognition toward systems that can better understand, reason about, and interact with the physical world. At the same time, the industry is learning how to bring these richer capabilities into devices constrained by power, cost, and size. Those two themes—more capable multimodal intelligence and the challenge of making it practical at the edge—are central to the 2026 Embedded Vision Summit.

To see why this matters, it helps to remember how far the field has come. Before deep neural networks became practical, computer vision algorithms were designed by hand: Engineers reasoned through a problem, wrote algorithms specific to that use case, and spent large amounts of time tuning and retuning them when conditions changed. Yann LeCun’s 2014 Embedded Vision Summit keynote helped crystallize a profound change: Instead of explicitly programming every rule, we could train general-purpose models by showing them examples. That shift transformed both what computer vision systems could do and the economics of developing them.

Now we are in the midst of another discontinuity. Vision-language models (VLMs) and related multimodal architectures are opening the door to a new class of products and systems: ones that combine images and language to become more adaptable, more intuitive, and more capable in unstructured environments. But they also raise practical questions for product developers: Can they run in embedded systems rather than only in the cloud? If they can, what kinds of hardware and software are required? These are exactly the kinds of questions we address at the Summit.

One of the most exciting topics this year is what may come after VLMs: “world models” and other more powerful types of models for physical AI. Our Day 1 keynote speaker, Eric Xing, president of the Mohamed bin Zayed University of Artificial Intelligence and a professor of Computer Science at Carnegie Mellon University, will discuss breakthroughs in world models. This work points toward AI systems that do more than interpret visual inputs or answer questions about them. Instead, they begin to model how the world behaves—an important step for robotics, autonomy, and all kinds of systems that must anticipate change and act effectively in dynamic environments. This theme appears throughout the conference in sessions on world models, robotics, autonomy, and the sensing and compute demands of increasingly capable machines.

Partner Content
View All
Is Copper Sintering the Key to Advancing Wide Band Gap Semiconductors?
Is Copper Sintering the Key to Advancing Wide Band Gap Semiconductors?
By Rejoy Surendran, Market Strategy Manager & Xinpei Cao, Sr. Principal, Application Engineering, Henkel 04.27.2026
SK hynix Receives 2026 IEEE Corporate Innovation Award
SK hynix Receives 2026 IEEE Corporate Innovation Award
By SK hynix 04.26.2026
Intelligent Cost Optimization is a Competitive Advantage
Intelligent Cost Optimization is a Competitive Advantage
By TechInsights 04.24.2026
If one side of the story is increasing capability, the other is implementation reality. Our Day 2 keynote speaker, Vikas Chandra, senior director at Meta Reality Labs, will address this in “Scaling Down Is the New Scaling Up.” For many developers, the main challenge is not simply creating a powerful model, but fitting sophisticated AI into products that must operate with tight constraints on energy, latency, bandwidth, and memory. Smart glasses and similar systems illustrate this transition well: To create compelling user experiences, these products need sophisticated perception and reasoning without always depending on a distant cloud service. Making that possible requires advances not just in models, but in architectures, compression, deployment tools, and system design.

That is why this year’s program emphasizes practical deployment so strongly. The Embedded Vision Summit has always aimed to occupy the sweet spot between academic research and marketing hype: The place where engineers and product teams can learn what actually works, what the tradeoffs are, and how to avoid common pitfalls. The 2026 program continues that approach with sessions spanning fundamentals, technical insights, business insights, and enabling technologies, plus an exhibit floor showcasing commercially available processors, sensors, cameras, tools, and software. Attendees can learn not only about emerging techniques, but also about the building blocks they can use now to create real products.

Deployment brings with it a host of challenges, and once edge AI and vision systems move beyond prototypes, the hard problems often shift away from model accuracy alone. Developers must handle fleet management, updates, data drift, hardware variation, and changing real-world conditions. For example, managing large fleets of deployed systems and incorporating feedback from the field to improve them over time is becoming increasingly important. To that end, we have Summit sessions on robotics at fleet scale, production deployment lessons, industrial pilot failures, and scaling vision-based retail systems from lab environments to thousands of real-world endpoints.

For attendees who want hands-on learning, the Summit will also offer training classes on vision-language models on Wednesday, May 13, beginning with an introductory session and continuing with a more advanced session focused on video understanding and agentic AI. That reflects another long-standing principle of the event: Concepts can be introduced in a talk, but deeper understanding comes from working directly with the technology; “you have to get your hands dirty.”

Fifteen years ago, many product creators still thought of computer vision as expensive, specialized, and impractical for mainstream products. Today, embedded AI and computer vision are being applied across retail, healthcare, agriculture, transportation, robotics, and wearables. The frontier has shifted. The question is no longer whether these technologies matter, but how quickly innovators can learn to use them effectively. The 2026 Embedded Vision Summit, taking place May 11-13 in Silicon Valley, is designed to help them do exactly that.