Scientist builds Age of Empires goat neural network to challenge claims of AI empathy

How a neural network built with virtual goats exposes the psychological illusion of empathy in modern artificial intelligence.

June 17, 2026

Scientist builds Age of Empires goat neural network to challenge claims of AI empathy
A whimsical yet highly methodical experiment has emerged from the intersection of retro gaming and cutting-edge computer science, delivering a sharp critique of how the modern artificial intelligence industry evaluates its own creations[1]. An AI scientist affiliated with Microsoft and the University of York recently constructed a fully functional, trained neural network inside the scenario editor of the 1999 strategy video game Age of Empires II[1][2]. What appears on the surface to be a humorous exercise in engineering is, in reality, a pointed academic attack on the widespread tendency to attribute human-like characteristics to large language models[1][2]. The research argues that the conversational, text-based interfaces of modern chatbots create a powerful psychological illusion[2][3]. By transferring the identical mathematical principles of a neural network to a medieval-themed virtual environment populated by digital goats, the underlying calculations remain unchanged, but the perceived intelligence completely vanishes[2][4].
To ground this critique in physical proof, the researcher first demonstrated that the computational engine of Age of Empires II is functionally and Turing-complete, meaning it is mathematically capable of performing any calculation that a modern supercomputer can[2][5][6]. Leveraging the game’s built-in scripting and scenario-editing tools, he designed a rudimentary neural network—specifically a one-bit perceptron—along with its corresponding training circuit[2][4]. Instead of running on traditional silicon microchips with electrical signals, this network operates using the pathfinding movements of virtual animals[2][4]. Goats act as the carriers of binary signals: a goat standing on grass represents a value of zero, while a goat positioned on a bridge represents a value of one[1][4]. The network's logical gates, constructed from a combination of XNOR and AND gates, utilize game triggers to process these positions[4]. To manage concurrency and prevent race conditions where signals arrive out of order, the developer constructed gate ready rails made of ice paths, which regulate the flow of the signal-goats[4][7]. When a calculation is completed, the original input goats are removed, and new output goats are spawned on their designated paths, proving that neural computations can be physically executed through the most absurd of mediums[4][7].
This bizarre setup illustrates a core concept in the philosophy of computation known as substrate non-uniqueness[8]. If a neural network can be run within a video game engine, it implies that any sufficiently complex substrate—whether it is Lego bricks, water pipes, or a highly coordinated postal service across a metropolitan area—can execute the exact same algorithms that power modern generative artificial intelligence[9][4][8]. The researcher poses a thought experiment: if one were to copy a large language model's architecture into the Age of Empires II engine and feed it the input I feel lonely, the underlying mathematics would eventually cause the goats to trigger a sequence of actions that spells out an empathetic response, such as I am sorry to hear that, perhaps you should talk to a friend[10][3]. While a user chatting with a web-based bot might easily believe the system genuinely understands loneliness, an observer watching virtual goats run across medieval bridges would never attribute empathy or consciousness to the game[2][3]. This disconnect reveals that the human-like traits frequently attributed to artificial intelligence are not inherent to the models themselves, but are instead projected onto them by humans who are seduced by natural-sounding chat interfaces[2][3].
To prove that this issue of anthropomorphism is a systemic problem within the scientific community rather than a fringe phenomenon, the researcher conducted a comprehensive review of the academic literature[11]. He compiled a dataset of 315 AI research papers published over a two-year period, spanning from mid-2024 to mid-2026[11][12]. Collected from major databases like Semantic Scholar and arXiv, and curated using advanced language model filtering, the papers were meticulously analyzed for how they framed and measured model behavior[11][13]. The findings were striking: 57 percent of the analyzed papers began with the pre-existing assumption that large language models possess human-like traits in their very premises[11][12]. Out of these, 36 percent went on to conclude that these traits were indeed present[11][12]. More alarming still was the analysis of the 47 papers that made human-like attributes the explicit subject of their research[11]. Among this subset, 77 percent concluded in favor of the models possessing anthropomorphic qualities, such as empathy, moral reasoning, anxiety, or self-awareness[11][14]. This high rate of positive findings points to a severe methodological vulnerability where researchers set up experiments that are highly prone to circular reasoning, ultimately validating the very assumptions they started with[14][15].
The implications of this study strike at the heart of the current artificial intelligence boom, offering a vital reality check to an industry prone to hype and over-attribution[16][3]. The research serves as a stark counterweight to high-profile incidents where engineers and public figures have claimed that models have achieved consciousness or shown signs of independent thought[11][17]. To correct this course, the author advocates for the universal adoption of what he terms the Null Assumption[2][18]. Under this methodological framework, AI evaluators must operate under the default assumption of substrate non-uniqueness, assuming that a model possesses absolutely no human-like psychological traits until rigorous, substrate-independent evidence proves otherwise[2][18][19]. By demonstrating that the most advanced AI architectures are functionally equivalent to goats wandering over ice ramps and wooden bridges, this research demands that the industry move away from romanticized, qualitative descriptions of machine mind and return to the objective, mathematical reality of computation[2][4].

Sources
Share this article