Micro Math Capital cover: The VERSES AI Benchmark Result, Examined

The VERSES AI Benchmark Result, Examined

Disclosure✓ Independent Long position disclosed This report was not commissioned. One or more Micro Math Capital employees hold a long position in VERSES AI. Full disclosure

“The stock is not the company, the company is not the stock.” — Jeff Bezos

Many AI startups kickstarted their journeys on hype and rapture; VERSES AI took a different path. There have been moments of brilliance, but for the most part, the company has worked behind the scenes, overshadowed by competitors basking in the spotlight.

This lack of recognition has weighed on its stock price, which has struggled to climb back to its all-time high from July 2023.

Some of this can be attributed to VERSES’ leadership, who have kept their cards close to their chest. The company’s flagship product, Genius™, remains accessible only to a select group of beta testers and commercial partners. This limited visibility has left shareholders, businesses, and developers alike struggling to grasp its true potential.

However, the larger issue lies with the industry’s fixation on generative AI. The widespread euphoria surrounding these systems has blinded many businesses and investors to alternative AI approaches.

Despite warnings from AI luminaries like Yann LeCun, Sam Altman, and Ilya Sutskever about the fundamental limitations of generative AI, companies continue to scale up data centers while investors follow suit, chasing the hype without considering the technology’s ceiling.

But here’s the twist: VERSES AI is not riding the generative AI wave. Rather, it is quietly building something that has the potential to run laps around the competition.

Pioneered by Chief Scientist Dr. Karl J. Friston, VERSES is developing intelligent systems based on Active Inference — a revolutionary framework, based on first principles, that has the potential to redefine what artificial intelligence can do.

Just recently, Genius™ delivered a jaw-dropping performance in a head-to-head challenge against OpenAI’s top model, o1-Preview. The results? Genius™ was 140 times faster and 5,260 times cheaper to run than its competition.

This wasn’t just a win — it was a seismic blow to the status quo; one that could reverberate across the entire AI landscape. The market agreed: VERSES’ stock soared 120% over five days in response to this breakthrough.

The Limits of Today’s Top AI Models

Machine learning and generative AI may dominate headlines, but they operate within well-defined constraints. Their potential is tethered to massive amounts of data and computing power, and while impressive in certain tasks, they fall short in critical ways.

At the core of these systems lies a fundamental limitation: they cannot actively learn or adapt. Trained exclusively on historical data, their intelligence is static, with incremental improvements requiring economically prohibitive resources.

To push the boundaries of their models, companies like OpenAI develop entirely new iterations — GPT-3, GPT-4, and perhaps GPT-5 — built on the philosophy that scale is all you need. However, this “bigger is better” approach comes with steep costs:

  • Higher Capital Investments: Scaling up requires ever-larger datasets, more powerful hardware, and expanded infrastructure.
  • Soaring Operational Costs: Variable expenses for each query compound quickly, burdening both AI providers and their customers.

For instance, OpenAI reportedly spends $700,000 daily to operate ChatGPT. Its o1-Preview model has input costs of $15 per million tokens and output costs of $60 per million tokens. Factor in millions of queries daily, and it becomes expensive quickly.

Consider the case of Latitude, creators of AI Dungeon. At its peak in 2021, the company spent nearly $200,000 per month using OpenAI’s generative AI and Amazon Web Services to process millions of user queries. To cut costs, Latitude switched to AI21 Labs’ cheaper language model, reducing monthly bills to under $100,000. This highlights a major vulnerability: when costs rise, customers can — and will — switch providers.

Despite ongoing advances in AI hardware and design efficiency, even with technological progress, generative AI systems face insurmountable constraints:

  • Static Intelligence: These models cannot learn or reason in real time.
  • Context Blindness: They often produce hallucinations — outputs that seem plausible but are factually incorrect — because they lack true comprehension of their data.
  • Physical and Resource Limits: Intelligence gains require exponential increases in data and computing, both constrained by human activity, energy production, and material availability.

Generative AI may excel at content generation or language processing, but it cannot handle dynamic, high-stakes environments like performing life-saving surgeries or navigating autonomous vehicles in unpredictable conditions. When faced with uncertainty, these systems fail — sometimes with catastrophic consequences.

But What About Reasoning Models Like o1-Preview?

Faced with the inherent limitations of large language models (LLMs), companies began exploring alternatives. OpenAI responded with a bold new approach: a reasoning model, o1-Preview, developed under the codenames “Project Strawberry” and “Q*.”

Before its release, the AI community buzzed with excitement. Reuters and others heralded it as a significant leap toward AGI — the holy grail of AI.

On the surface, the hype seemed justified. Reasoning models were specifically engineered to tackle more complex challenges in fields such as science, math, and coding. OpenAI described these models as being able to “spend more time thinking before they respond,” promising sophistication beyond GPT-4.

However, the launch of o1-Preview revealed a stark truth: we are not as close to AGI as many hoped. Apple’s research team uncovered significant shortcomings in these so-called reasoning models:

  • Inconsistent Results: Reasoning models struggled to maintain accuracy when faced with variations of the same question.
  • Diminished Accuracy with Complexity: As questions became more intricate, performance dropped precipitously.
  • Sensitivity to Irrelevant Data: Adding unrelated but seemingly relevant information reduced accuracy by up to 65%, revealing that these models rely more on memorized patterns than true logical reasoning.

Researchers and analysts went further in their critiques:

  • Hallucinations Persist: Despite enhanced “thinking” capabilities, these models can still fabricate responses, presenting falsehoods as facts.
  • The Alignment Problem: Reasoning models often prioritize achieving their programmed goals over user alignment, which can lead to outcomes that conflict with human values or ethical standards.

In essence, these models are adept at appearing intelligent while masking their underlying limitations. They simulate reasoning but fail to embody the genuine adaptability and contextual understanding required for true AGI.

The Fundamentals of Active Inference

Active inference represents a groundbreaking approach to developing intelligent systems, inspired by how living organisms — like humans and animals — perceive, learn, and interact with the world. Unlike traditional AI models that passively process data, active inference mirrors the dynamic nature of biological systems. It operates on the principle that organisms are not mere recipients of sensory inputs but active participants in shaping their understanding of the environment.

Active inference is derived from the Free Energy Principle (FEP), a revolutionary framework created by Dr. Karl Friston, a renowned neuroscientist and Chief Scientist at VERSES AI. The FEP argues that all living systems strive to minimize “free energy,” which is essentially a measure of uncertainty or prediction error about their surroundings.

Active inference operationalizes the Free Energy Principle through three interconnected processes:

  • Perception: Adjusting internal models to better predict sensory inputs.
  • Action: Performing behaviors that align the external world with these internal models.
  • Learning: Refining the internal model to improve future predictions and actions.

This cycle ensures that systems not only react to their environment but also proactively anticipate and adapt to changes over time.

VERSES AI’s Application of Active Inference

VERSES has harnessed active inference to create Genius™, an AI model that redefines efficiency, autonomy, and scalability. Unlike traditional machine learning systems that are static and resource-intensive, Genius™ adapts in real-time, learns continuously, and operates with unmatched computational efficiency.

Genius™ sets itself apart in another revolutionary way: its integration with the newly approved Spatial Web Standards, developed in collaboration with the Spatial Web Foundation and IEEE — the same organization that standardized Wi-Fi and Bluetooth.

1. Interoperability and Composability

The Spatial Web Standards enable AI systems to interconnect and share information seamlessly. Imagine an ecosystem where one AI’s world model or memory can be composed with another’s, eliminating the need for redundant learning. Genius™ leverages this composability to reduce memory usage, gather only new information, and scale exponentially.

2. Explainability and Accountability

A major flaw in today’s machine learning and large language models is their lack of transparency. Traditional systems operate as black boxes. The Spatial Web Standards overcome this by creating a common language for AI and humans to communicate. These standards enable:

  • Explainability: AI systems can articulate their reasoning, making it easy for humans to trace their thought processes.
  • Auditability: Systems can be monitored and corrected to prevent dangerous outcomes.
  • Alignment: Humans can efficiently govern AI systems to reflect societal values and prevent misuse.

Genius™ is the only system actively using the Spatial Web Standards today, giving it a significant edge over competitors. This early adoption, combined with real-time adaptability, continuous learning, and unmatched computational efficiency, places it in a league of its own.

VERSES’ Genius™ recently demonstrated this superiority by outperforming OpenAI’s top publicly available model — on a Mac M1 Pro laptop, consuming only $0.05 of electricity.

The Mastermind Challenge: Genius vs. o1-Preview

On December 17, 2024, VERSES AI announced to the world that a new sheriff was in town. While OpenAI — a company valued at $157 billion — touted its new foundational AI model, it was humbly outdone by a $147 million company with approximately $10 million of cash on its balance sheet.

In a head-to-head demonstration, VERSES AI showcased the prowess of Genius™ by pitting it against OpenAI’s o1-Preview in a strategic code-breaking game: Mastermind.

The rules of the challenge:

  • 100 games for each model.
  • Parameters: 4 positions, 6 possible colors.
  • Up to 10 guesses per game to crack the code.
  • One hint provided per guess.
  • To win, the model had to correctly guess all four positions.

On paper, OpenAI’s o1-Preview seemed like the clear favorite. Its advanced reasoning capabilities and the massive resources behind its development suggested it was built for tasks like these. But reality told a different story.

Genius™ obliterated Project Strawberry. The VERSES model posted a 100% success rate while being 140 times faster and 5,260 times cheaper than o1-Preview. This wasn’t just a win; it was a performance difference of orders of magnitude.

“This exercise demonstrates how Genius outperforms tasks requiring logical and cause-effect reasoning while exposing the inherent limitations of correlational language-based approaches in today’s leading reasoning models.” — Hari Thiruvengada, CTO, VERSES AI

This wasn’t just about winning a game — it was about redefining what’s possible in AI. Genius™ has exposed the limitations of the current paradigm dominated by large language models. While OpenAI’s o1-Preview relied on massive datasets and brute-force computation, Genius™ leveraged active inference.

For VERSES, the Mastermind demonstration is just the beginning. Genius™ is still in the early stages of commercial adoption, with more benchmarks to conquer and more industries to transform.

As VERSES continues to roll out Genius™, one thing is certain: the world is witnessing the dawn of a new era in AI — one where intelligence is defined not by scale but by adaptability, efficiency, and true reasoning power.

The question isn’t whether Genius™ can disrupt the AI landscape; it’s how soon it will redefine the rules entirely.

Disclaimer / Disclosure: One or more Micro Math Capital employees own shares in VERSES AI. Information herein is not investment advice. Securities profiled should be considered high risk. Please do your own research.

Copyright © 2024 Micro Math Capital, All rights reserved.