I wonder how much better an LLM could be given even better training data.
For example, the total number of tokens contained in the physical and digital collections of a moderately-sized university library is (probably) equal to or on par with the size of the training data for GPT 3.5.
What would happen if you could train just on that? I know we're using huge training sets, but how much of it is just junk from the internet?
(There should be some representative junk in the dataset, but nowhere near the majority.)
For example, the total number of tokens contained in the physical and digital collections of a moderately-sized university library is (probably) equal to or on par with the size of the training data for GPT 3.5.
What would happen if you could train just on that? I know we're using huge training sets, but how much of it is just junk from the internet?
(There should be some representative junk in the dataset, but nowhere near the majority.)