From Frontier AI Research to Building CreativAI’s “SQL Layer” for the Physical World
For most of his career, Mohamed Elhoseiny (Founder & CEO of CreativeAI) lived at the frontier of artificial intelligence research. His journey took him through KAUST, Stanford, Meta FAIR, and Adobe, giving him a rare view of AI from both sides: as a researcher pushing the boundaries of what machines could understand and as an engineer watching those technologies move into products used by billions of people. But somewhere along that journey, a different question began to emerge—one that would eventually lead him away from research papers and toward entrepreneurship.
What happens after a machine learns to see? For Elhoseiny, that question became the foundation for CreativAI, the company he is building with Waleed Shaarani, Founding CCO. Their ambition is not simply to make computer vision models more accurate. They want to create an entirely new layer of infrastructure that allows organizations, AI agents, and robots to program against what machines see. The founders describe it simply: CreativAI is building the SQL layer for Visual and Physical AI.
A Career at the Frontier of Visual AI
Elhoseiny’s path to CreativAI was shaped by years spent working at the intersection of academic research and commercial AI. At Meta FAIR, he experienced what it meant for visual intelligence to operate at extraordinary scale. Computer vision was no longer simply an academic discipline; it was becoming infrastructure for products used by billions of people, influencing everything from how content is understood to how users experience platforms such as Instagram.
His research work also explored a fundamental shift in how machines interact with images. Projects including VisualGPT and MiniGPT-4 investigated whether AI could genuinely understand and reason about visual information rather than simply classify individual pixels or objects. Those experiments pushed forward an important frontier in multimodal AI, but they also exposed a much larger challenge. The question was no longer whether a machine could understand an image. The question was what humans could actually do with that understanding at scale.
The Question That Became CreativAI
That distinction became increasingly important as AI moved from understanding individual pieces of content toward interacting with the physical world. If a machine can perceive a factory floor, a hospital, an airport, a warehouse, or a road, how can engineers actually build applications on top of that perception?
Elhoseiny saw a parallel in the history of computing. Before SQL, vast quantities of structured information existed inside databases, but interacting with that information required fundamentally different methods. SQL created an abstraction layer that allowed developers and businesses to query data without rebuilding the underlying system for every question.
CreativAI is attempting to create a similar abstraction for visual information. The platform continuously transforms what cameras observe into a structured knowledge base containing people, objects, actions, and events. Instead of treating video as a collection of files that someone must manually search through, the founders want organizations to be able to query their physical environment as naturally as they query a database. That shift is what they mean by “SQL for Visual and Physical AI.”
Why Bigger AI Models Won’t Solve Everything
The obvious response to this problem is to build better AI models. Over the past few years, the industry has demonstrated how rapidly larger models and larger datasets can improve machine perception. But Elhoseiny believes there are structural limitations that cannot simply be solved by scaling.
The first is the lack of a universal way to structure visual data. More than 80% of internet traffic is visual, and the world now has roughly a billion cameras and counting, yet visual information remains fundamentally different from text or tabular data in how businesses interact with it. Today, answering a new question about video can require a new model, pipeline, and engineering project.
The second problem is even more difficult: the long tail of real-world events. An enterprise may care deeply about an unusual operational event that barely exists in any public training dataset. A general-purpose model may simply never have seen enough examples of that event to recognize it reliably. Making the model larger does not automatically solve the problem because the missing ingredient is not necessarily model capability—it is knowledge of that organization’s specific environment.
CreativAI’s approach is to move the problem from building another model to creating structured, adaptable data. If a company wants to ask a new question, the goal is for that to become a query rather than a months-long machine-learning project.
From Video Footage to a Living Knowledge Base
This is where CreativAI’s vision begins to diverge from conventional computer vision. Traditional systems can detect objects, identify people, or generate descriptions of what appears in a video. CreativAI instead aims to create structured representations of entities, events, and behaviors that can be continuously queried and updated. That distinction matters because structured information can become part of an organization’s operational systems. A video clip might tell someone what happened at a particular moment. A structured record can allow an enterprise to ask what happened across thousands of locations, identify patterns, compare events, and allow AI agents or robots to act on the resulting information.
There is another critical element: traceability.
CreativAI’s system links generated records back to the exact moment in the original footage from which they were derived. The objective is therefore not simply to make visual information searchable, but to make it queryable, traceable, and verifiable. That creates a different kind of relationship between AI and enterprise data. Instead of asking, “Can we find this video?”, an organization can begin asking, “What happened across our entire operation, and what were the exceptions?”
For Elhoseiny and Shaarani, that is the beginning of a much larger transformation.
Making Perception Programmable
The founders believe that once perception becomes programmable, the implications extend far beyond video analytics. People, cameras, AI agents, and robots could potentially operate from a shared understanding of the same physical environment. A robot could query what has happened around it. An AI agent could reason over the same information. A human operator could investigate an event and trace it directly back to the underlying footage. Instead of multiple systems maintaining disconnected versions of reality, CreativAI envisions a shared layer of structured information. The founders call this ambient intelligence—a world where the physical environment is continuously understood by machines and that understanding becomes available to everything operating within it.
It is an ambitious proposition, but one that reflects how dramatically the relationship between AI and the physical world is changing. The next challenge is no longer simply teaching machines to see. It is teaching the world to become data that machines can understand, query, and act upon. And that is the problem Mohamed Elhoseiny and Waleed Shaarani have set out to solve with CreativAI.
How CreativAI Is Turning the Physical World Into Queryable Data
The promise of CreativAI becomes clearer when the founders move from the abstract idea of “visual intelligence” to the environments where it can create measurable value. For Mohamed Elhoseiny and Waleed Shaarani, the biggest opportunity is not another application that sits on top of computer vision. It is the infrastructure underneath the growing number of cameras, AI agents, autonomous systems and robots operating in the physical world.
Today, a camera can see an event without necessarily creating usable knowledge from it. A robot can perceive its surroundings but may lack persistent memory of what happened before. An enterprise can have thousands of cameras recording continuously while still relying on humans to search through footage after something has already gone wrong. CreativAI’s proposition is to turn those fragmented observations into structured information that can be continuously queried, understood and acted upon.
Where Physical AI Meets Real-World Operations
Robotics and autonomous vehicles are among the most immediate applications because machines operating independently cannot rely on humans to interpret every video frame. A person can watch a clip and understand what happened; a machine needs structured information that it can reason over and use in the next decision. CreativAI is designed to provide that persistent visual memory, allowing physical AI systems to work with information that accumulates rather than starting from scratch with every new frame.
But the opportunity extends well beyond robots. Consider an airport with thousands of cameras, baggage systems, gates, security infrastructure and ground operations. Each system may observe a different part of the environment, yet those observations often remain disconnected. The same challenge appears in factories, hospitals and logistics networks, where important events occur continuously but the ability to convert those events into shared operational intelligence remains limited.
The founders describe this as “a robot’s problem without a robot’s architecture.” As enterprises increasingly deploy AI agents, those agents will need the same kind of persistent, structured visual memory that autonomous machines already require. CreativAI’s strategy is therefore to build the infrastructure layer before the market fully realizes that it needs one.
Beyond Factories and Robots
The visual-data opportunity also extends into industries that may not immediately be associated with Physical AI. Media companies, broadcasters, sports organizations and content libraries sit on enormous archives of video, much of which remains difficult to search based on what actually happens inside the footage. Once that information becomes structured, decades of previously underutilized content could become searchable and usable in entirely new ways.
That same principle applies to public safety, healthcare and other environments where visual events matter but cannot always be manually monitored. The value is not simply identifying objects or producing descriptions; it is building a continuously updated representation of what is happening, allowing organizations to ask questions across their operations rather than investigating individual clips one at a time.
This distinction is central to CreativAI’s thesis. Search finds something you already know to look for. Structured intelligence can reveal something you did not know to ask.
Solving the Long Tail of Visual Intelligence
One of the company’s most significant technical ambitions is addressing what Elhoseiny describes as the “pre-training gap.” General-purpose AI models can be remarkably capable, but they cannot be expected to have encountered every rare event that matters inside a particular factory, hospital, airport or autonomous vehicle operation. The most commercially important events may sometimes be precisely those that occur too infrequently to feature meaningfully in general training datasets.
CreativAI’s approach is to allow organizations to adapt the system to their own visual environments and define the events, entities and behaviors that matter to them. Instead of treating every new requirement as a model-training exercise, the company wants users to be able to work with visual intelligence more like a programmable data layer. That makes the system capable of evolving as an organization’s operations change and new questions emerge.
The architecture also emphasizes traceability. When CreativAI generates structured information, each record can be connected back to the precise moment in the source footage that produced it. For high-assurance environments, that distinction matters: an organization should not have to accept an AI-generated summary without being able to investigate the evidence behind it.
From Research to Enterprise Deployment
Turning an ambitious research concept into enterprise infrastructure requires a different kind of discipline. Elhoseiny acknowledges that the hardest challenge has not necessarily been whether the technology works, but deciding where to focus first. CreativAI’s underlying engine can potentially serve robotics, hospitals, factories, airports, media archives and public safety. That breadth creates an unusual problem for an early-stage company: almost every conversation can reveal a legitimate use case. The temptation is to pursue them all simultaneously, but the founders believe that doing so would create several partially developed products rather than one category-defining platform.
Their answer is sequencing. The company intends to identify the largest opportunity, go deep enough to win that market, and then move into the next vertical while reusing the same underlying technology. The advantage is that the core engine does not need to be rebuilt for every industry; the questions and domain-specific requirements change, while the fundamental infrastructure remains consistent. That approach could ultimately allow CreativAI to expand across industries without becoming a collection of unrelated vertical products.
Building Toward a Shared Intelligence Layer
The larger vision is a world where humans, AI agents and robots no longer operate with separate understandings of the same physical environment. Today, a robot may have one perception system, a security platform another, and an operations team a third. Each creates its own interpretation of reality, and humans are often left to reconcile the differences.
CreativAI wants to provide a shared structured record that all three can use. A robot could act on it, an AI agent could reason over it, and a human could query it in natural language—all while tracing the information back to its original visual evidence. The objective is what the founders describe as moving from a collection of disconnected systems toward “one body, one mind.”
The scale of the opportunity is being reinforced by the rapid growth of Physical AI and robotics. Elhoseiny points to the significant capital flowing into robotics as an indicator of where the broader market is heading, while noting that long-term market forecasts should be treated as directional rather than definitive. His more important observation is that every new autonomous machine creates more demand for infrastructure capable of understanding and reasoning over the visual world.
For CreativAI, therefore, the market opportunity is not simply today’s spending on video analytics. It is the much larger possibility of making the physical economy itself programmable. And that sets the stage for the company’s next challenge: proving that an idea born from frontier AI research can become the infrastructure enterprises actually deploy—and eventually, infrastructure they no longer think twice about using.
From Frontier Research to Global AI Infrastructure
For Mohamed Elhoseiny and Waleed Shaarani, building CreativAI is ultimately about something larger than computer vision. The founders are trying to establish a new category of infrastructure for a world in which cameras, AI agents, autonomous systems, and robots increasingly interact with the same physical environments. If successful, CreativAI could become the layer that allows these systems to share a common, structured understanding of what is happening around them.
That ambition is already shaping how the company approaches customers, partnerships, product development, and its next stage of growth. Rather than attempting to become another application competing for attention in an increasingly crowded AI market, CreativAI is positioning itself further down the stack—where visual intelligence becomes infrastructure that other products and organizations can build upon.
A Founder-Led Approach to the Enterprise Market
CreativAI’s go-to-market strategy is deliberately hands-on. At this stage, the founders are working closely with enterprises and design partners rather than attempting to scale through a purely self-service API model. The objective is to understand complex operational environments directly, deploy the technology inside them, and use those deployments to refine both the product and the company’s understanding of where its infrastructure creates the greatest value.
That approach also creates a potential distribution advantage. Some customers may eventually integrate CreativAI beneath their own products and services, allowing the platform to reach larger numbers of end users without requiring the startup to build a separate sales organization for every industry. For a company developing horizontal infrastructure, that combination of direct enterprise relationships and strategic distribution could become increasingly important as the platform expands.
The company is currently working with design partners in areas including autonomous vehicles and robotics, while also seeing applications across visual-data infrastructure, public safety, media archives, and industrial operations. Those early relationships are important not only as potential commercial opportunities but also as real-world environments in which CreativAI can test the boundaries of its technology.
Building With the Infrastructure Giants
CreativAI’s development has also benefited from support from some of the world’s largest technology companies. The startup has received backing through programs including AWS, Google for Startups, Microsoft for Startups, and NVIDIA Inception, giving its team access to infrastructure, compute, technical resources, and broader ecosystem support.
For an AI infrastructure company, access to compute can materially affect the pace at which a technical team can experiment and iterate. The founders have particularly highlighted the role of AWS and Google in providing early infrastructure and compute support, allowing the company to direct more of its resources toward research and engineering rather than carrying the full cost of infrastructure from day one.
But infrastructure alone does not create a company. The more difficult task is turning technical capability into something enterprises can reliably deploy, adapt, and depend upon. That is where the combination of Elhoseiny’s research background and Shaarani’s commercial leadership becomes particularly important.
The Challenge of Choosing Where to Win
One of CreativAI’s biggest strategic challenges is also a consequence of its technology’s breadth. The same underlying platform could potentially address problems in robotics, healthcare, factories, airports, logistics, media, and public safety. Almost every enterprise conversation can therefore reveal a legitimate use case.
For an early-stage company, that can become a dangerous form of abundance. Elhoseiny describes the discipline required to resist pursuing every opportunity simultaneously. The company intends to identify the largest opportunity, go deep enough to win it, and only then expand into the next vertical. Because the underlying engine remains reusable, the founders see sequencing—not reinvention—as the path to building a horizontal platform without becoming distracted by too many partially developed products.
That philosophy reflects a broader lesson in building AI infrastructure: technological possibility is not the same as commercial focus. A platform can theoretically serve dozens of markets, but a startup still has to decide where it can establish undeniable product-market fit first.
A Market Being Created in Real Time
The founders believe the opportunity around Physical AI is expanding alongside the rapid growth of robotics and autonomous systems. Elhoseiny points to the substantial capital flowing into robotics as an important indicator: every new generation of autonomous machines creates more visual information and, consequently, greater demand for systems capable of understanding that information.
He argues that CreativAI should not be measured simply against today’s video analytics software market. That market reflects what organizations currently spend on tools for working with video. The larger opportunity, in his view, is the value created when the physical economy itself becomes instrumented, structured, and queryable.
This distinction is central to the company’s long-term thesis. If robots, autonomous vehicles, smart facilities, and AI agents become increasingly common, visual intelligence will become less of a standalone feature and more of a foundational requirement. The infrastructure that allows these systems to share information could ultimately become as fundamental to Physical AI as databases became to modern software.
The Five-Year Vision
CreativAI’s five-year ambition is therefore not simply to become a successful AI application company. The founders want the company’s technology to become something organizations assume will exist whenever they deploy intelligent systems into the physical world.
That is an ambitious comparison. Databases were once specialized infrastructure; today, almost every serious software system assumes that some form of persistent data layer exists underneath it. CreativAI is betting that visual intelligence could follow a similar trajectory as physical environments become increasingly connected to AI. The implications extend beyond enterprise efficiency. A shared visual intelligence layer could allow humans, AI agents, and robots to operate from the same source of truth. A person could ask a question in natural language, an AI agent could reason over the answer, and a robot could act on the same structured information—all while retaining a connection to the original evidence.
That is the larger idea behind CreativAI: not simply teaching machines to see, but creating the infrastructure through which seeing becomes understanding, understanding becomes data, and data becomes action. For Elhoseiny and Shaarani, the opportunity is to build that infrastructure before the physical world becomes fully intelligent. And if their vision plays out, CreativAI may eventually become the kind of technology that users rarely see—but that countless AI systems, robots, and enterprises quietly depend upon every day.


