Fable 5's Vision-Only Pokémon Sapphire Run: A New Benchmark
An AI agent, codenamed Fable 5, has successfully completed a vision-only run of Pokémon Sapphire, a feat that not only confirms its advanced capabilities but also sets a new benchmark for AI in complex, sequential decision-making tasks. This specific run, conducted on August 7, 2026, utilized a Chinese fan translation of the game and was completed in a single, uninterrupted session lasting 7 hours and 11 minutes. The agent was capped at 2,000 turns, where each turn represented a single screenshot and a single decision, translating to an API call.
The primary objectives of this run were not to discover novel strategies but to validate existing hypotheses about Fable 5's emergent behaviors. The experiment focused on three key areas: first, how Fable 5 would construct its own internal manual after its 'notes rule' was reset to a blank page; second, how it would approach button presses after its 'press-count advice' was removed; and third, to assess the completeness of its memory regarding the Hoenn region (Generation III) compared to its established knowledge of the Kanto region (Generation I, as tested in a previous FireRed run).
The conclusions drawn from the run were affirmative. Fable 5 successfully generated its own operational guidelines, adapted its interaction strategy without explicit press-count advice, and demonstrated a robust memory of the Hoenn region. Perhaps most striking, the agent played the game faster than its previous run of Pokémon FireRed, which was itself a demonstration of advanced AI gameplay. This suggests a significant optimization in Fable 5's processing or decision-making architecture when faced with familiar yet distinct game environments.

The Vision-Only Paradigm: How Fable 5 Perceives the World
The 'vision-only' condition is crucial here. Unlike agents that might have direct access to game state information (like Pokémon stats, inventory, or map data), Fable 5 operates solely on the visual input it receives. This means it interprets the game world through screenshots, much like a human player would. This constraint forces the AI to develop sophisticated visual processing and state-tracking mechanisms. For this run, the traditional 'notes' that an AI might keep—acting as an external memory or state tracker—were deliberately reset to a blank page. This compelled Fable 5 to rely more heavily on its internal memory and its ability to infer game state from visual cues alone.
Furthermore, the removal of 'press-count advice' meant that the AI could not rely on pre-programmed heuristics for how many times to press a button (e.g., 'press A twice to open the menu'). It had to learn or infer the appropriate number of presses based on visual feedback or learned patterns. This tests the agent's adaptability and its capacity for fine-grained interaction with the game's interface.
The comparison to the FireRed run is particularly telling. Pokémon FireRed, being an enhanced remake of the original Red/Blue, represents a foundational experience in the Pokémon franchise. Fable 5's ability to not only match but exceed the speed of that previous run in Sapphire indicates a significant leap in its ability to generalize and optimize its performance across similar, yet distinct, game worlds. It suggests that the underlying architecture is not just memorizing specific game paths but is developing a more abstract understanding of game mechanics and player objectives.
Implications for AI Development and Gaming
This experiment provides compelling evidence for the efficacy of large language models (LLMs) like Fable 5 in complex interactive environments. The ability to process visual input, maintain long-term memory, and make sequential decisions over thousands of turns is critical for applications ranging from advanced game AI to robotics and autonomous systems. The speed at which Fable 5 navigated Pokémon Sapphire, especially under the constrained vision-only and blank-notes conditions, suggests that current LLM architectures are becoming increasingly capable of handling real-time, dynamic challenges.
The specific confirmation that Fable 5's memory of Hoenn is as complete as its memory of Kanto is a subtle but important finding. It implies that the AI's knowledge base is not simply a static repository but is actively managed and updated, allowing for detailed recall across different game generations. This has profound implications for how we think about AI's capacity for learning and retention in complex, simulated or real-world scenarios.
What remains to be explored is the precise nature of the 'manual' Fable 5 wrote for itself. Understanding these emergent internal strategies could unlock new methods for AI training and control. If an AI can autonomously generate effective operational guidelines, it suggests a path towards more self-sufficient and adaptable artificial agents. The speed increase over the FireRed run also warrants deeper analysis; was it due to optimized pathfinding, more efficient decision trees, or a better understanding of the specific mechanics of Sapphire?
The success of this vision-only run underscores the potential of AI agents to engage with and master complex systems using only sensory input, mirroring human learning processes more closely than ever before. It challenges the notion that AI requires direct state access to perform at a high level, pushing the boundaries of what's possible in artificial general intelligence research.
