Introducing Bart: A Linguistic Time Capsule
Unbounded Labs has unveiled Bart, a novel large language model (LLM) with 2.82 billion parameters. What sets Bart apart is its training data: 20.1 billion tokens of English text exclusively sourced from before 1931. This deliberate choice positions Bart not as a competitor in the race for the most advanced general-purpose AI, but as a specialized tool for linguistic historical research and a unique exploration of how language and thought evolved in a different era. The project, which consumed three months and approximately $800 in compute costs, aims to uncover the subtle nuances, biases, and stylistic conventions embedded within early 20th-century and pre-20th-century English.
The decision to train an LLM on such a constrained and historical dataset is a departure from the current industry trend, which favors massive, diverse, and contemporary datasets. Bart's creators at Unbounded Labs are not chasing state-of-the-art performance on modern benchmarks. Instead, they are focused on the qualitative aspects of language. By limiting the training corpus to text predating the significant linguistic shifts of the mid-20th century, Bart offers a unique lens through which to examine the evolution of grammar, vocabulary, sentiment, and societal perspectives as captured in written form.

Why a 'Vintage' LLM?
The motivation behind creating Bart stems from a desire to understand the foundational elements of language and communication before the advent of modern media, globalized communication, and the digital age. LLMs trained on contemporary data often reflect the vast, complex, and sometimes chaotic information landscape of the present day. Bart, by contrast, is designed to act as a linguistic time capsule. Its responses and generated text are expected to embody the style, tone, and underlying assumptions prevalent in early 20th-century literature, journalism, and personal writings.
This approach allows for several distinct applications. Researchers can use Bart to analyze historical texts, identify patterns in language use over time, or even generate text that mimics a specific historical style for creative or academic purposes. For instance, a historian studying the early 1900s might use Bart to explore how common topics were discussed, what kinds of metaphors were prevalent, or how social attitudes were implicitly conveyed through language. The model could also serve as a fascinating counterpoint to modern LLMs, highlighting the dramatic shifts in linguistic expression and worldviews over the past century.
The choice of 1931 as a cutoff date is significant. This period predates many of the major global events that reshaped societies and, consequently, language—including World War II, the rise of mass media in its current form, and the rapid technological advancements of the latter half of the 20th century. English language usage in 1931 would have been influenced by Victorian and Edwardian literary traditions, early industrialization, and a very different set of social and cultural norms.
Technical Details and Accessibility
Bart is built upon a 2.82 billion parameter architecture, a size that is substantial yet considerably smaller than many of the frontier models currently dominating the LLM landscape. This scale is sufficient to capture complex linguistic patterns without requiring the astronomical computational resources of models with hundreds of billions or trillions of parameters. The training process involved 20.1 billion tokens, carefully curated from a corpus of English written before 1931. This meticulous data selection is key to Bart's unique character.
Unbounded Labs has made Bart accessible through several channels. A live demo is available at unboundedlab.com/chat/bartholomew, allowing users to interact with the model directly. For those interested in the technical underpinnings and development process, an article detailing the project is available at unboundedlab.com/blog/bartholomew. Furthermore, the model weights are available on Hugging Face at huggingface.co/jbduran/bartholomew-sft, enabling developers and researchers to download, experiment with, and build upon Bart.
The project's modest cost of $800 and three-month development timeline suggest an agile and focused approach. This efficiency, combined with the specific niche of historical linguistics, indicates a potential strategy for smaller labs or research groups to contribute meaningfully to the LLM space without needing the vast capital reserves of major tech corporations. It highlights that innovation in LLMs can come not only from scaling up but also from creative data curation and targeted research questions.
Broader Implications for LLM Development
The existence of Bart prompts a reconsideration of the LLM development paradigm. While the pursuit of larger, more general models continues, there is a clear value in specialized LLMs trained on curated datasets. Bart demonstrates that a focused approach can yield unique insights and tools. It raises the question: what other historical or specialized linguistic corpora could be leveraged to create LLMs that offer distinct perspectives or capabilities? Consider an LLM trained solely on Shakespearean plays, or one trained on legal documents from the 18th century. Each would offer a unique window into a specific domain of language and thought.
Furthermore, Bart's existence highlights the importance of understanding the data that underpins AI models. LLMs are not neutral observers; they are reflections of their training data, complete with the biases, assumptions, and cultural contexts of the information they consume. By deliberately training on pre-1931 text, Unbounded Labs is not just creating a language model; they are creating a tool that can help us interrogate the historical underpinnings of our current linguistic structures and societal views. What biases are implicitly encoded in language from a century ago, and how do they compare to those embedded in today's LLMs? Bart provides a tangible way to explore these questions.
The success of Bart, even on a small scale, could encourage more experimentation with niche datasets and specialized LLM architectures. It suggests that the future of AI might not be a monolithic landscape of giant, general-purpose models, but a diverse ecosystem of specialized tools, each designed for a specific purpose or to explore a particular facet of knowledge or language. For developers, this means new opportunities to work with and fine-tune models that cater to unique historical or domain-specific needs, moving beyond the one-size-fits-all approach.
