Reece Shuttleworth

(pre-)training a language model

Training loss over pretraining Evaluation loss over pretraining

During this project, I implemented several transformer models and variants along with training code from scratch in PyTorch. My largest training run was a 2B model trained on ~100B tokens on a single 8xH200 node rented on AWS (credits courtesy of YC). This model, while nowhere close to the ability of current models, was still able to learn strong language understanding.

Below is an example of the completion of the trained model to the prompt "The capital of France is". We see a correct answer ("Paris") and clear English language. However, we see the completion soon turn unfactual (Paris is not "the second largest in the world") and repetitive (a common observation in early LLMs).

Input: The capital of France is
Output:  Paris. It is the largest city in Europe and the second largest in the world. Paris is located in the north of France, on the Seine River. The city is the capital of the Paris Region, which is the largest in France. The city is also the capital of the Paris Region, which is the largest in France. The city is also the capital of the Paris Region, which is the largest in France. The city is also the capital of the Paris Region, which is the largest in France. The city is also the capital of the Paris Region, which is the largest in France. The city is also the capital of (...)