TL;DR: I tried to make an LLM play Wordle. This is the postmortem, because I decided to use GRPO instead of SFT. Using GRPO, I was able to get a 10% win rate with RL. With SFT, that shot up to 50%. Ultimately, it cost a bunch of days wasted and $10 of GPU compute wasted.
Definitions
Setting up the environment
Hopefully, you know what Wordle is, but if you don't: it's a game where there's a random five-letter word and you get six guesses to figure it out. Each guess reveals, per letter, whether it's green, yellow, or gray, and you use that to refine your future guesses until you finally get it.
I built a training environment that simulates Wordle. Here's how it looks:
The goal
Now that we have the environment, let's define the goal. Our goal is to make a solver that maximizes win rate and minimizes guesses.
Wordle is technically solvable deterministically; 3Blue1Brown has a very good video on that.
The goal of this project was to get an LLM to basically do the same thing. For an LLM, there were two ways we could take the output:
- Greedy: take the single most-likely guess.
- Sampled: draw many different LLM outputs, then post-process them to pick a guess.
This matters a lot for small models. With greedy, the most likely guess might only be right 30% of the time, and there might be four other very good guesses that each show up 10% of the time. Also, small models might not have great distributions yet; the winning word might not have trained enough to become the most likely, but might still be one of the candidates in there.
The Plan
For this task, I wanted to try out the Pi harness, just for fun. Usually I use OpenCode, but I was making big changes to my developer environment, so I decided now is a great time to try Pi.
For the model, I used DeepSeek V4 Flash, because I am poor and have no money :(
Since machine learning tasks take a very long time, I got the pi-goal extension, so Pi could do long-running tasks and work towards things, just like Codex does.
The plan, in order:
- Built an entropy-based solver. A scripted solver that picks the guess maximizing expected info gain. The solver that I created had a 100% win rate with 3.46 average turns on the full answer list.
- SFT warmstart from solver traces + GRPO. This was probably a mistake. I don't know why we decided to do GRPO here, but it was recommended by DeepSeek and I thought it was interesting to try it out to see if we could get good results from this.
- Reward v1: partial reward for greens and yellows, +1 for winning, penalties for invalid replies (no guess, or a non-word): win=+1.0, +0.02/green, +0.01/yellow, −0.1 invalid, −0.1 repeat.
Problems
I trained on this for a bit before I realized there were a few problems with the plan.
- SFT had too little data. The warmstart was 3,000 solver-generated games, using LoRA r16 for a quick tuning pass on Qwen3-0.6B. For each task, the LLM was given the Wordle board and then asked to make a guess for the Wordle game. It was able to pick up some things, but not enough to get a good training pass. After the SFT, the greedy eval was approximately 0%, so we were unable to tell if SFT taught anything usable.
- The model had a bad win rate after GRPO. 7.5% win rate greedy, sampled ~30% win rate.
I decided to look into the GRPO part first. Looking at my reward, I discovered it wasn't great. Out of all the pool of possible guesses, a lot of guesses may get a reward of 0 (all greys), which doesn't lend itself well to hill-climbing. Thus, we changed the reward to model information gain based on how much a guess narrows down the pool of potential guesses:
reward = log2(|candidates_before| / |candidates_after|) + win bonusAfter doing a run with this, the results came in: greedy 10%, sampled ~30%. This was better, but still pretty bad.
Additionally, I also forgot to keep track of the number of steps used for both GRPO runs, which means I couldn't tell if we had enough steps to learn a pattern, or if we ran into a different problem such as overfitting.
Going with SFT
At this point, I decided to go with SFT as the primary path and avoid GRPO for now. The realization behind it: a reward that says "reduce the candidate set" and an expert solver's trajectory are the same supervision; the solver already is the optimal reply to that reward. GRPO was re-deriving, through gradient descent, what a solver trace states directly.
The first problem to tackle is the 3,000-game training set: it's too little to get a good SFT training run.
The 3,000 games came from a restricted list of 500 answer words, which had the solver play each one from six different openings, with a bit of noise mixed in so the traces weren't all identical.
To increase the training set, I made these three changes:
- Split the games into per-state rows (game state so far -> next guess): 3,000 games -> 160,371 rows (train 111,961 / val 24,043 / test 24,367).
- Diversify setup of previous guesses. We originally used the solver, which would pick ideal guesses. By inserting non-optimal priors, we can diversify the game states so that the model can learn from games with random or weak guesses.
- Improve reward data with top-k labels. Instead of labeling each state with only the solver's single best guess, we saved the top-k best candidate guesses as multiple valid labels. Top-k labels give the model more information about what guesses are good instead of teaching the model to make a singular specific guess.
Improving the LLM harness
The LLM sometimes makes a guess that violates the model's own constraints — reusing a gray letter, moving a green off its spot, or putting a yellow in a ruled-out position. Since the model can see the game state, the most-likely token should never contradict it.
I made a harness that fixes this by sampling 8 guesses, filtering to ones consistent with all feedback, and picking the most-voted survivor.
Results
The SFT-only result: greedy 45-52%, harness (sampled + verified) 77-80%. This was much better than the GRPO version, so I decided to stop here — mostly because of how much I'd spent on RunPod instances, and how much time I'd put into training the model.
What I could have done better
I think this was a good experience in trying GRPO and seeing where it would work and where it wouldn't. I also learned a lot about SFT from this experiment. If I were to do this again, I would not use GRPO, since this isn't a good use case for it.
I would also work on improving my training set and keeping track of the number of steps used. That would make it easier to make good decisions based on my previous runs, and to figure out what the model is doing right and what it still needs to work on.
Conclusion
Overall, this was a nice learning exercise into fine-tuning LLMs to achieve a simple task. I was able to use SFT to get the LLM to be sufficiently capable when paired with a surrounding harness for playing Wordle.