Training GPT-2 on Nvidia DGX Spark
*This is a summary post of my experiments taking an existing GPT harness & architecture and training it on custom datasets found online.
After getting my DGX Spark, I wanted to test its capabilities by training a GPT-2 model. I used the nanoGPT framework by Andrej Karpathy, which is a lightweight and efficient implementation of GPT-2.
I’m going to walk you through the steps I took to set up the training environment, prepare the dataset, and run the training process on the DGX Spark.
Setting Up the Environment
I used Andrej Karpathy’s nanoGPT framework, which in his Reproducing GPT-2 video, explains clearly how to set up the script and optimizers to train his GPT-2 replica. I largely followed his instructions but with some minor tweaks to fit the size of model I wanted to train. I used PyTorch and the HuggingFace Transformers library to handle the model architecture and training loop. I also made sure to install the necessary dependencies, including CUDA and cuDNN in my uv environment, to leverage the GPU capabilities of the DGX Spark.
Dataset Preparation
I asked ChatGPT for datasets to try. Karpathy in his GPT-2 training video used the Tiny Shakespeare dataset, which is a small dataset of Shakespeare’s works. My attempts at trying to find more interesting datasets brought me to the TinyStories and the WikiText datasets. Both are available on HuggingFace as a free download and contain a variety of text data that can be used for training language models. I downloaded both datasets onto my local disk from HuggingFace, created .bin files, and imported them during training as necessary.
I do recommend using as creative of datasets as possible, but for the sake of this specific experiment, it doesn’t matter as much. The goal is to get output that resembles regular speech rather then complete gibberish.
Experiment 1: Overfitting a Batch
In my first experiment, I wanted to see if I could overfit a single batch of data. The reason is that if the model can overfit a single batch, it means that the model is capable of learning and memorizing. I used the Tinystories dataset split into training, validation, and test splits. I limited the test token number to 20,000 and the validation / test splits to 5000 tokens. The script that I used to run the training can be found here: prepare.py
For this run, I ran using the following hyperparameters:
model
n_layer = 2 n_head = 4 n_embd = 128 block_size = 128
training
batch_size = 16 learning_rate = 1e-3 max_iters = 2000
dropout
dropout = 0.0
The model wasn’t too big and the batch size was small, so I expected it to quickly overfit. Overfitting is more of a smoke test to see if the model is capable of learning and memorizing. I ran this on my DGX Spark and after 2.5 hours, it achieved a loss of 2.3536.
To generate samples, I used the following command:
python sample.py \
--out_dir=out \
--dataset=tinystories10m \
--num_samples=5 \
--max_new_tokens=200
The sample file (linked here) contains code to sample from the model checkpoint that is generated at the end of every successful trial run.
This is what it gave back to me:
fter that, the big dog came to the park. The dog loved to play and play in the park. The dog saw Sue and Sue under the bed. Sue was scared, but she did not help.
Sue closed the bushes and said, "I want to play with you!" They played and laughed and had fun. When the dog got to Tom, they decided to play together. Sue was a nice girl. They played a lotion with Sue, and they played together.<|endoftext|>Once upon a time, there was a girl named Lucy. She had a friend named Jerry. Tom loved to play with Sue and Sue. They had fun computer all day long. They would laugh and enjoy their dance.
One day, Tom's mom came to play with all the little ones. But Tom hurt and did not mind. Tom was sad. He thought, "I am, Tom. I love to feel better."
Tom tried to use the computer, but he could not help. Tom did not want to play with Lily. So, Tom tried to balance on the computer. He was not happy to take the computer with him. He tried to push the computer, but he was difficult. Tom was sad, and he did not want to hurt Sara.
Tom tried to climb the computer, but he hurt his knee. He fell down and started to worry. The computer did not move. The computer made Tim's face and came back.
Sara and Tom were sad. They did not want to go. They did not know what to do. They said they were sorry and lost. They had a bad ending.<|endoftext|>Once upon a time, there was a little dog named Spot. Spot lived in a small house with many toys. One day, Tim went to the store with his mom. They were on the street.
Tim's mom saw grown-up, and said, "Lily, can we have enough food with this crayons!" They looked at each other and saw a big, red bucket. They were curious and excited.
"Look, the box is in the grass!" Tim said. He wanted to pick the box and open it, but it was difficult to get it. So, he started to break it. He pulled and pulled, and the box.
"Oh no, what are you doing?" Lily asked.
Lily said, "We need to cut the box, Lily."
"Give it
It looked pretty good, but the story didn’t make much sense, and there were end of text tokens sprinkled in here and there. Regardless, the model achieved a very low loss and was able to largely memorize the batch.
Experiment 2: Training on TinyStories Dataset
For the next experiment, I wanted to see if I could train the model on a more larger dataset without truncating any number of tokens. I decided to use the TinyStories dataset in full, which is about 10 million tokens.
I divided the data into train, val, and test splits just like before.
I changed the hyperparameters to the following:
model
n_layer = 12 n_head = 12 n_embd = 768 block_size = 1024
training
batch_size = 12 learning_rate = 6e-4 max_iters = 2000
For the sampling, I also added the following parameters:
sample_max_new_tokens = 100
sample_start = "\n"
sample_temperature = 0.8
sample_top_k = 200
metrics_log_file = 'metrics.jsonl' # structured, machine-parseable log (one JSON object per line)
train_log_file = 'train_baseline.log' # human-readable log, mirrors stdout
The final training run took about 4 hours, and it did come with a caveat*. The model was able to achieve a final loss of 0.24, and the loss generated looked like this:
*The training took a very long time, and it was because the evaluation split was doing multiple forward passes quite frequently. I ended up having to lower the number of passes the model made in order to speed up training.
The text result came back like
"No, Tom, I'm sorry. You were so silly and foolish. You can't take away the vase. You have to help me! I'll give you a big hug. You're my sister!"
Lila felt very relieved. She hugged Tom and Tom.
"Thank you, Tom. You're my best friend. You're my best friend. You're my best friends. You comfort me."
Tom smiled. He hugged Lila and Lila.
Lila hugged them back. She was glad they were safe. She was happy.
"I'm so glad you're okay, Lila. You're my best friends. You are more famous friends. You saved the vase and my love. You are very cool. You saved my heart and my heart. You are my best friends."
They hugged them back. They said goodbye to their stories. They played with their toys, their books, and their smile. They were happy. They were famous twins. They had a friend. They loved them. They were happy.<|endoftext|>Once upon a time, there was a lazy dog named Max. He loved to lazy and sleep all day. One day, Max was very tired and wanted to sleep.
Max's owner, a kind girl named She would give Max a warm bath. She would throw a soft blanket around Max's and hold it.
After Max felt better, he was not lazy anymore. He went to the park and played with his friends. They had a great time playing together.
The moral of the story is that if you try hard and stay lazy, or you might find a new friend.<|endoftext|>Once upon a time there was a little girl named Ann. She had a favourite dress, and she loved to bounce. One day, Ann decided to go on a sailing. She jumped up on her parents's shoulder and they began to spin around with a rhythm. Suddenly, Ann saw a sail by the sea. She wanted to find out what it was, so she decided to follow the sail.
Ann followed the sail for a long time, but she got closer and saw that it was actually a warning. She looked embarrassed, but she decided to stay and watch the sailboat instead. Suddenly, school panicked. Everything around her silently cheered.
Suddenly, a magical fairy appeared and told Ann that it was a magical sailingboat that Ann had not been seen in the sea, but she knew
The setences are more coherent and the story is tructured and kind of makes sense??!
Experiment 3: Training on WikiText Dataset
For the final experiment, I tested the model on the WikiText dataset, which contains about 103 million tokens. To prepare the data, I modified the prepare.py script to handle the larger dataset and split it into training, validation, and test sets.

I achieved a final loss of about 3.49 after 2000 iterations.
The generation came out to this for the 1024 block size:
<|endoftext|> Tōgō , under the tutelage of one of the most prominent Chinese characters in the series to a series written by the series . Although written by Tōgō , the story takes place over a period of four years .
<|endoftext|> Tōgō : Tōgō : The Story of Tōgō is a concept album for the series . The third installment , Tōgō : The Story of Tōgō , is a collaboration with series creator Yurok Miyagi , while Tōgō 's story is based on the series . The third installment , Tōgō : The Story of Tōgō , is the first to be written by Tōgō in which the character was originally shown . It was composed and produced by Tōgō . Each of the episodes to be played out of the series takes its name from the original to the original . The second installment , Tōgō : The Story of Tōgō : The Story of Tōgō , is the first to be written by Tōgō . The story of the series , however , is not considered to be an " modern romance " . The story of Tōgō is composed by Kazuichi Teng . The story arc is one of the most popular English @-@ language novels in Japan , reaching number 21 on the Sino @-@ Japanese language charts .
<|endoftext|> = = Plot = =
<|endoftext|> Tōgō : The Story of Tōgō : The Story of Tōgō is a story of the same name . In this story , the demon Zayō ( 杭 ; 毉绘 , " Oan " Tōgō ) is revealed to be from the series 's epigod , having been converted into a human form by the demon Zayō . The story follows Tōgō and the three other demons , a reincarnation of his tribe Tōgō , and a clan leader named Nōgō . The four men are identified as Nōgō 's clan and , in the story , are featured in the story . The story is set in the summer of the monthly evening .
<|endoftext|> The story of Tōgō is first presented in a different way ; the third installment was written by Zayō . The story is the only one in the story — the first one
I would say that it is far less impressive then the previous generations but with the 100 million token dataset, it would require more epochs of training in order to fully fi this to the correct distribution. I would say that the model is still learning and has not yet fully converged to the lowest possible loss.
Experiment 4: Training on Different Block Sizes
My next experiment was to see how the block size affected the training and generation. I ran the same training on the WikiStories dataset with different block sizes: 512, 256, and 128 tokens. I kept the other hyperparameters the same and ran the training for 2000 iterations each time.
Here are the final loss curves for each block size:
512 tokens: [INFO] iter 2000: loss 3.6946, time 6978.68ms, mfu 12.22%, tok/s 9,391

256 tokens: [INFO] iter 2000: loss 4.0265, time 6592.34ms, mfu 11.27%, tok/s 4,971

128 tokens: [INFO] iter 2000: loss 4.5207, time 1829.06ms, mfu 10.49%, tok/s 8,958

My next steps are to profile the attention layers to see which layers are holding things up. I anticipate that also using different variations of flash attention would speed up the calculations by not needing to materialize the intermediate matrices.
The entire repo can be found here:
CZ