Dataset of Five Centuries of Memory
Try to imagine a corridor of books stretching across five centuries. Stacked there are philosophical manuscripts, ancient maps, religious texts, literature, history, and even fable illustrations from a very distant time. Then a scanning machine arrives. A camera photographs each sheet. A computer reads the writing. Images are separated from the pages. The entire contents of the paper slowly turn into data. This means that today, a student in Jakarta or a researcher in Istanbul no longer needs to fly to London, England, to search for data. They simply open a laptop, type a keyword, and a portion of that old book appears on the screen. This giant project is not merely uploading files. It is the transfer of human memory into a machine.
It all began in 2005. The British Library in England partnered with Microsoft to scan tens of thousands of rare books published between 1510 and 1900. Microsoft poured millions of dollars into the project at the time. When the partnership concluded, the British Library Labs went a step further. In 2013, they cropped more than one million public-domain illustrations from those pages and released them for free on Flickr Commons. Now, that giant repository, a 626 GB trove containing 1.08 million images from 25 million book pages, has been repackaged. It currently resides on Hugging Face under the name British Library Book Images. The data is staggering. Imagine a file size of 626 Gigabytes. If you use an average smartphone with 128 Gigabytes of storage, that data alone would require nearly five full phones just to hold the images. That is not even counting the 25 million pages of text that must be read, parsed, and stored in computer memory. That figure far exceeds the stack of physical paper sheets that can fit inside an ordinary library building. This is a digital ocean of information, the product of thousands of years of human brains. The oldest book in the collection was published in 1510. That means there are traces of human thought over five centuries old.
The digitisation project tears down the walls of the building. Previously, a rare book had only one address: London. Now, the data from that book can move anywhere. The memory of five centuries has been transformed into raw data. A book is no longer just an object on a shelf. It has become data that can be searched, counted, compared, or concocted as material to train artificial intelligence (AI) systems. Simply put, a giant collection of data that has been grouped and organised so that computers can easily read it is called a dataset. If a regular book is read by humans page by page, a dataset is raw material consumed all at once by a computer. Inside it, thousands of texts, images, and supporting notes have been arranged so that a machine can process them without confusion. With this, researchers can count how many times a specific word appears across thousands of books. Linguists can simply track the shift in a word’s meaning across the ages. Historians can compare how a nation depicted cities, wars, and even diseases. The machine learns to recognise patterns from all of this.
But there is a major problem. This collection is not an honest mirror of all human history. It is a mirror that chooses its own angle. Of the 1.08 million images on Hugging Face, one-third come from the 1890s. Collections from before 1800 account for only 1.6 percent. If you use the entire dataset to train an AI without fixing the distribution, your AI will be obsessed with the Victorian era and understand nothing of other times. There is an even more fundamental issue. The machine did not scan every book that ever existed on earth. The machine only scanned books owned by a specific library, which happened to survive disasters, and were then selected by humans to be digitised. The machine never chose its own food. Humans decided. Whoever chooses which books to scan determines what knowledge will later be easily found by AI. A giant project like this is not unique to England. Google Books in the United States has already scanned more than 40 million books in 400 languages. For out-of-copyright books, Google allows you to download them for free. There is also HathiTrust, a collaboration of research universities that stores 17 million digital volumes containing 6.2 billion pages. Their system is even capable of analysing the text of millions of books at once.