The Japanese did not place much time or emphasis on encryption. So, if by help you mean "vacuumed up" all the stuff that was out in the open that was easy to crack, then yes.
Is it? How many different books are we talking about, and how much information is that, after conversion to text and lossless compression? Images, maybe, but text?
These models are trained on way more than just books. GPT-3 was trained on about half a terabyte of filtered plaintext and the training corpuses have grown significantly by then by all accounts.
They aren’t trained on compressed plaintext so I’m not sure of the relevance there. But regardless it’s my understanding that’s modern models are trained with orders of magnitude more storage than their parameters require. But it’s possible I’m incorrect. This is getting to the fringe of my knowledge of concrete LLM details.
The relevance is because the LLMs are storing information, not the explicit text, so we want to know how much actual information they need to store (this being an information theoretic argument). The representation in the parameters doesn't necessarily need 1 parameter per character, if the text is highly redundant.
A well-organized system derived from binder clips works for small/medium scale collections of debound books. It's a quick and easy, and relatively cheap, way to "rebind" them in parts. The clips even provide good surfaces for labels.
Archivists recommend standard ambient conditions or a little drier for long term storage. As you've said, too dry and the pages fall apart; permanent damage.
I suppose the fancy silica gels that maintain specific humidities would work in bags.
My understanding is that the fibers in paper can become brittle if they are too dry. However, this doesn't cause the paper to break down, it just makes it vulnerable to damage if flexed. So storing paper dry is fine if it is then humidified before being stressed.
It's possible. My only experience is with vintage newspapers that were allowed to dry out, and then when trying to open them just shatter like a carbonized volcano scroll.
It's truly humbling how much knowledge professionals from different fields have, and how easy it is to fall into a Dunning–Kruger effect trap where a bit of knowledge — extremely extrapolated — would've led me down the path of hubris and guessing so many wrong answers. Thanks all.
>> If I did that I'd seal them with silica gel to keep the humidity down.
> I would think this totally dries the pages out and then they just turn to dust, from experience.
They used to sell Boveda two-way humidity control packs that would keep a bag at a constant low-ish 32% humidity, but last I looked those were discontinued (it looks like they've pivoted pretty heavily to marijuana storage and higher-humidity products).
Yes, clearly it's completely absurd to think that shredding rare books to more cheaply train LLMs is anything but a moral good, which is why there is this entire thread is full of people jumping through hoops to explain how it's technically legal (and therefore fine) and "you wouldn't have bought those books anyway" (I suppose we don't have the choice anymore!) and, most amusingly, "they're actually becoming digitized and searchable this way" (are LLMs stochastic parrots or aren't they?).
Given that in the US we dispose of 640,000 tons of books each year, this attempt to inflate outrage is just ridiculous. Books are not things of intrinsic moral value; they are morally neutral physical objects.
To the extent the information in a book is something to preserve, this scanning is a good thing.
It happens on a massive scale, so it's completely reasonable. Books are disposable information delivery vehicles now. Publishers pulp great quantities of unsold books too.
These are increasingly available online, btw. Historical research is accelerated when historians have direct access to scans of relevant source material. Not destructively scanned, of course.
reply