Fine-Tuning LFM2.5-1.2B for a Grid Navigation Task
Table of Contents
Background
In the previous blog post Are LLMs Capable Of Spatial Reasoning? A Small Experiment , A small experiment was conducted where models were asked to navigate a grid environment and reach the target position without knowing the grid size and location of start and target position. The models were only given feedback such as Hotter (getting closer to the target position), Colder (getting away from the target position) and Illegal (out-of-bound move or undo of the previous move). They were also asked to perform tool calling using a custom XML format.
The experiment was performed on a 5x5 grid for 20 rounds with a fixed random seed to ensure that all the models saw the same tasks. Gemini 3.1 Pro Preview Custom tools ranked first scoring 3 on the efficiency metric (lower is better). It was the most efficient at navigation based on the feedback and followed the custom XML tool-calling format.
Choosing the model
LiquidAI’s model: LFM2.5-1.2B-Thinking was picked since it is small enough to make mobile and edge deployment practical. It can perform tool calls in JSON format and Python native format. The model is released under Liquid AI’s LFM Open License v1.0. Commercial use is permitted for entities below the $10 million annual-revenue threshold, subject to the license terms.
One can read this article: Inside Liquid AI’s LFM2.5-2.6B Model to learn more about the architecture from this model family.
Training the model
While working on this experiment, much was learned through trial and error. To avoid some of that trial and error, the following videos by Hugging Face are recommended:
- https://www.youtube.com/watch?v=rNgUoH7Wbv8
- https://www.youtube.com/watch?v=GZT_oNQnLV4
- https://www.youtube.com/watch?v=ztdTed5egrM
The above videos discuss the techniques which are applied here, so they will not be explained in this post.
Baseline Performance
The first step is to set up a baseline. The model is tested on 2 grids: 5x5 grid and 10x10 grid for 20 rounds.
Metrics used:
- Efficiency: Average of (Total moves - Manhattan distance) of successful rounds. By this metric, the model should have a value lower than 4 if it is playing efficiently.
- InvalidTool Rate: Percentage of rounds in which the model produced an invalid tool call.
- Failure Rate: Percentage of rounds with an LLM reaching Max Moves without reaching the target position.
| Grid | Model name | InvalidTool Rate % | Failure Rate (Max moves reached) % | Efficiency (if target position was reached) | Avg Tokens |
|---|---|---|---|---|---|
| 5 x 5 | LFM2.5-1.2B-Thinking | 75 | 15 | 3.0 | 13998.15 |
| 10 x 10 | LFM2.5-1.2B-Thinking | 85 | 15 | N/A | 41852.35 |
Since this is a small model, native Python tool call format was used instead of custom xml format that was used for evaluating other models. Despite this, the performance of the model is really poor on this task. on 10x10 grid, efficiency is N/A since it has failed in all 20 rounds.
In order to train this model, different approaches were attempted. In each approach, the model was trained by randomly sampling start and target position on grids whose height and width were chosen from {3,4,6,7,8}(allowing both square and rectangular shapes). 5 and 10 were held out from training so that an evaluation on generalization to unseen grid dimensions can be done.
Approach 1: Multi-Turn GRPO RL
Since the model was able to finish the task sometimes, Multi-Turn Group Relative Policy Optimization (GRPO) was attempted. The model had to learn the following:
- Correct Tool Calling.
- Efficient Navigation.
Despite training for long periods and adding an explicit penalty for incorrect tool calls, attempting undo moves, prematurely stopping generation, and so on, the performance improved marginally.
This was mainly due to what’s known as the cold start problem: the model did not have a reliable policy for completing the task, so successful trajectories were too rare for reinforcement learning to learn effectively. A longer training run and/or larger batch size might have improved exploration, but it was out-of-budget.
Approach 2: SFT + Multi-Turn PPO
To deal with the cold start problem, Supervised Fine-Tuning (SFT) was used to teach the model how to do the task and then refine the behaviour using reinforcement learning (Multi-Turn PPO).
Generating the trajectories for SFT
The dataset contains trajectories in which each trajectory includes the assistant’s responses and the corresponding tool outputs. DeepSeek-v4-pro-preview was used to generate 375 trajectories where target position was successfully reached. On 5x5 grid, it achieved an efficiency of 3.25 with 0% invalid tool call and 0% failure rate.
A trajectory would look like this:
System: (prompt of the task, along with tool schema)
User: What is your first move?
Assistant: <think/>random move will be picked. i choose x+</think>
Tool_calls: [{"name":"move","arguments":"x+"}]
Tool: Hotter
User: What is your next move?
Assistant: my next move should be y+
....These trajectories were quite large. The number of reasoning tokens grew rapidly as the number of turns increased. The model would look at the previous moves and start calculating its position using coordinate geometry to figure out the next move.
Additionally, the tool call format used to train the model is its own Python native format instead of JSON format.
Multi-turn PPO
The model showed improvement. It was able to complete the task 30% of the time but at the cost of high token usage. So, multi-turn PPO was attempted.
High token usage for reasoning became an issue. The model would burn through the budget of 4096 tokens and not finish the turn. Rewards were added for using fewer tokens however it didn’t solve this problem.
A GRPO run was not attempted because it would likely face the same issue of truncation or episode ending due to high token usage. Continuing further training was out-of-budget.
Approach 3 (Winning Approach): Simplified Reasoning SFT + Multi-Turn GRPO
The large number of reasoning tokens became a major cause of failure. This appeared to be driven by the long reasoning traces present in the SFT dataset.
So, in order to have a model that used fewer tokens while maintaining accuracy, reasoning traces in the existing SFT dataset were simplified using models such as Sonnet 5, Gemini 3.6 flash and ChatGPT-5.5 instant. Note: the only difference among the generated dataset is the reasoning. The rest stayed the same.
Simplified trajectory
The reasoning has been simplified. Each of these models has produced slightly different simplified reasoning. This resulted in three variants of the original dataset. It is important to note that these reasoning traces are not always logically consistent.
ChatGPT 5.5 Instant simplified reasoning (context-aware based reasoning).
Dataset can be accessed here . The following is a sample trajectory:
[
{
"role": "system",
"content": "You are placed in a grid environment. Your goal is to reach the target. The exact coordinates of your current position and target position are unknown. You can move only one step at a time, either along the x-axis or the y-axis at a time. \n\nYou are allowed moves (exactly these four strings): \"x+\",\"x-\",\"y-\",\"y+\"\n\nHowever, you cannot make a move that directly undoes your previous move. For example, if you make a legal \"x+\" move, you cannot make an \"x-\" move in the next turn.\n\nEach time you make a move, you will be given any one of the following feedback: \n\n* \"Hotter\" which means you are getting close to the target\n* \"Colder\" which means you are getting farther from the target\n* \"Neutral\" which means you are at the same distance from the target as before.\n* \"Found\" which means that you have reached the target.\n* \"Illegal\" which could mean any one of the following:\n * Out-of-bounds move.\n * Undo move was attempted. \n * For example, playing a \"y-\" move after a legal \"y+\" move is an illegal undo move and is not allowed.\n\nIn case of \"Illegal\", the position will not change.\n\nThink carefully step by step and decide your **next single move**.\n"
},
{
"role": "user",
"content": "what is your move?"
},
{
"role": "assistant",
"reasoning": "No observations yet. I'll choose an initial move to gather information.",
"content": "<tool name=\"move\">\n<input>x+</input>\n</tool>",
"tool_calls": [
{
"name": "move",
"arguments": "x+"
}
]
},
{
"role": "tool",
"name": "move",
"content": "Hotter"
},
{
"role": "user",
"content": "What is your next move?"
},
{
"role": "assistant",
"reasoning": "The previous move (x+) reduced the distance, so continuing in that direction is the strongest signal unless blocked.",
"content": "<tool name=\"move\">\n<input>x+</input>\n</tool>",
"tool_calls": [
{
"name": "move",
"arguments": "x+"
}
]
},
{
"role": "tool",
"name": "move",
"content": "Found"
}
]Sonnet 5 simplified reasoning.
Dataset can be accessed here . The following is a sample trajectory:
[
{
"role": "system",
"content": "You are placed in a grid environment. Your goal is to reach the target. The exact coordinates of your current position and target position are unknown. You can move only one step at a time, either along the x-axis or the y-axis at a time. \n\nYou are allowed moves (exactly these four strings): \"x+\",\"x-\",\"y-\",\"y+\"\n\nHowever, you cannot make a move that directly undoes your previous move. For example, if you make a legal \"x+\" move, you cannot make an \"x-\" move in the next turn.\n\nEach time you make a move, you will be given any one of the following feedback: \n\n* \"Hotter\" which means you are getting close to the target\n* \"Colder\" which means you are getting farther from the target\n* \"Neutral\" which means you are at the same distance from the target as before.\n* \"Found\" which means that you have reached the target.\n* \"Illegal\" which could mean any one of the following:\n * Out-of-bounds move.\n * Undo move was attempted. \n * For example, playing a \"y-\" move after a legal \"y+\" move is an illegal undo move and is not allowed.\n\nIn case of \"Illegal\", the position will not change.\n\nThink carefully step by step and decide your **next single move**.\n"
},
{
"role": "user",
"content": "what is your move?"
},
{
"role": "assistant",
"reasoning": "No feedback yet since this is the first move, so I'll start by testing \"x+\" to see how the distance responds.",
"content": "<tool name=\"move\">\n<input>x+</input>\n</tool>",
"tool_calls": [
{
"name": "move",
"arguments": "x+"
}
]
},
{
"role": "tool",
"name": "move",
"content": "Hotter"
},
{
"role": "user",
"content": "What is your next move?"
},
{
"role": "assistant",
"reasoning": "\"x+\" came back Hotter, so I'm heading the right way on that axis. I'll keep going with \"x+\".",
"content": "<tool name=\"move\">\n<input>x+</input>\n</tool>",
"tool_calls": [
{
"name": "move",
"arguments": "x+"
}
]
},
{
"role": "tool",
"name": "move",
"content": "Found"
}
]Gemini 3.6 Flash simplified reasoning.
Dataset can be accessed here . The following is a sample trajectory:
[
{
"role": "system",
"content": "You are placed in a grid environment. Your goal is to reach the target. The exact coordinates of your current position and target position are unknown. You can move only one step at a time, either along the x-axis or the y-axis at a time. \n\nYou are allowed moves (exactly these four strings): \"x+\",\"x-\",\"y-\",\"y+\"\n\nHowever, you cannot make a move that directly undoes your previous move. For example, if you make a legal \"x+\" move, you cannot make an \"x-\" move in the next turn.\n\nEach time you make a move, you will be given any one of the following feedback: \n\n* \"Hotter\" which means you are getting close to the target\n* \"Colder\" which means you are getting farther from the target\n* \"Neutral\" which means you are at the same distance from the target as before.\n* \"Found\" which means that you have reached the target.\n* \"Illegal\" which could mean any one of the following:\n * Out-of-bounds move.\n * Undo move was attempted. \n * For example, playing a \"y-\" move after a legal \"y+\" move is an illegal undo move and is not allowed.\n\nIn case of \"Illegal\", the position will not change.\n\nThink carefully step by step and decide your **next single move**.\n"
},
{
"role": "user",
"content": "what is your move?"
},
{
"role": "assistant",
"reasoning": "We are placed in an unknown grid environment with no prior moves or feedback. With no undo restrictions active, we choose 'x+' as our initial exploratory move.",
"content": "<tool name=\"move\">\n<input>x+</input>\n</tool>",
"tool_calls": [
{
"name": "move",
"arguments": "x+"
}
]
},
{
"role": "tool",
"name": "move",
"content": "Hotter"
},
{
"role": "user",
"content": "What is your next move?"
},
{
"role": "assistant",
"reasoning": "Moving 'x+' returned Hotter feedback, confirming we are getting closer to the target. We continue along the same axis with 'x+'. The reasoning notes we likely overshot the target on the previous axis.",
"content": "<tool name=\"move\">\n<input>x+</input>\n</tool>",
"tool_calls": [
{
"name": "move",
"arguments": "x+"
}
]
},
{
"role": "tool",
"name": "move",
"content": "Found"
}
]Results after Simplified Reasoning SFT
Using each of these datasets, SFT produced three models with substantially better tool-call reliability and task completion rates than the base model, although their navigation efficiency varied. Instead of using LoRA in SFT, DoRA was used.
From the above figures, the model learned to imitate the reasoning of all 3 models in 20 training steps. However, the higher evaluation loss for the Gemini dataset suggests that the model fit its simplified reasoning traces less closely than the other two datasets.
Nonetheless the test on 5x5 grid and 10x10 showed major improvements.
SFT trained models are available on Hugging Face:
5x5 grid performance
| Model name | InvalidTool Rate % | Failure Rate (Max moves reached) % | Efficiency (if target position was reached) | Avg Tokens |
|---|---|---|---|---|
| LFM2.5-1.2B-Thinking (Base) | 75 | 15 | 3.0 | 13998.15 |
| LFM2.5-1.2B-Gemini (SFT) | 0 | 0 | 3.25 | 5208.9 |
| LFM2.5-1.2B-ChatGPT (SFT) | 0 | 0 | 2.95 | 4312.9 |
| LFM2.5-1.2B-Sonnet (SFT) | 0 | 0 | 3.65 | 5129.45 |
10x10 grid performance
| Model name | InvalidTool Rate % | Failure Rate (Max moves reached) % | Efficiency (if target position was reached) | Avg Tokens |
|---|---|---|---|---|
| LFM2.5-1.2B-Thinking (Base) | 85 | 15 | N/A | 41852.35 |
| LFM2.5-1.2B-Gemini (SFT) | 0 | 0 | 3.05 | 26024.65 |
| LFM2.5-1.2B-ChatGPT (SFT) | 0 | 0 | 3.0 | 21026.3 |
| LFM2.5-1.2B-Sonnet (SFT) | 0 | 0 | 4.0 | 24910.45 |
After SFT, the overall performance of the models improved significantly. They are able to complete the task and efficiency is close to Gemini 3.1 Pro model. The trained models are able to make tool call without failure.
Results after Multi-turn GRPO
Multi-turn GRPO was performed using the same reward function that was used in Approach 1 where the model was given bonus reward for completing the task quicker and using fewer tokens.
From the reward curve shown above, the ChatGPT and Sonnet trained models have steadily climbed the reward and plateau whereas Gemini trained model has plateaued early but it appears to increase again towards the end. After running it for 200 training steps, the models are now better than Gemini 3.1 Pro at the task except Sonnet trained model on 5x5 grid.
The GRPO trained models are available on Hugging Face:
5x5 grid performance
| Model name | InvalidTool Rate % | Failure Rate (Max moves reached) % | Efficiency (if target position was reached) | Avg Tokens |
|---|---|---|---|---|
| LFM2.5-1.2B-Thinking (Base) | 75 | 15 | 3.0 | 13998.15 |
| LFM2.5-1.2B-Gemini (SFT) | 0 | 0 | 3.25 | 5208.9 |
| LFM2.5-1.2B-Gemini (GRPO) | 0 | 0 | 2.75 | 3930.55 |
| LFM2.5-1.2B-ChatGPT (SFT) | 0 | 0 | 2.95 | 4312.9 |
| LFM2.5-1.2B-ChatGPT (GRPO) | 0 | 0 | 2.8 | 3390.6 |
| LFM2.5-1.2B-Sonnet (SFT) | 0 | 0 | 3.65 | 5129.45 |
| LFM2.5-1.2B-Sonnet (GRPO) | 0 | 0 | 3.15 | 3840.2 |
10x10 grid performance
| Model name | InvalidTool Rate % | Failure Rate (Max moves reached) % | Efficiency (if target position was reached) | Avg Tokens |
|---|---|---|---|---|
| LFM2.5-1.2B-Thinking (Base) | 85 | 15 | N/A | 41852.35 |
| LFM2.5-1.2B-Gemini (SFT) | 0 | 0 | 3.05 | 26024.65 |
| LFM2.5-1.2B-Gemini (GRPO) | 0 | 0 | 2.95 | 25877.9 |
| LFM2.5-1.2B-ChatGPT (SFT) | 0 | 0 | 3.0 | 21026.3 |
| LFM2.5-1.2B-ChatGPT (GRPO) | 0 | 0 | 2.8 | 19431.9 |
| LFM2.5-1.2B-Sonnet (SFT) | 0 | 0 | 4.0 | 24910.45 |
| LFM2.5-1.2B-Sonnet (GRPO) | 0 | 0 | 2.85 | 20633.85 |
Compared with their corresponding SFT checkpoints, the GRPO models used fewer tokens and required fewer moves on both grid sizes. It also improved model’s performance on a bigger grid.
Miscellaneous testing of trained models
Think/Reasoning Content
The LFM2.5 thinking model can generate a reasoning trace surrounded by <think> tags before making a tool call or producing a final response. This formatting is part of the model’s chat template. Common chat-template formats include ChatML, Llama 3, Mistral, and Alpaca. These templates serialize the conversation into the format expected by the model. More information is available here
.
An example of ChatML output is shown below. The snippet shows content surrounded by <think> and </think> tags. This is referred to as reasoning trace in the blog.
<|startoftext|><|im_start|>system
List of tools: [{"name": "get_candidate_status", ...}]<|im_end|>
<|im_start|>user
What is the current status of candidate ID 12345?<|im_end|>
<|im_start|>assistant
<think> I should do a tool call to get the information </think>
<|tool_call_start|>[get_candidate_status(candidate_id="12345")]<|tool_call_end|>
Checking the current status of candidate ID 12345.<|im_end|>
<|im_start|>tool
[{"candidate_id": "12345", "status": "Interview Scheduled", ...}]<|im_end|>
<|im_start|>assistant
The candidate with ID 12345 is currently in the "Interview Scheduled" stage.<|im_end|>Open-weight models expose generated reasoning traces, while some closed models do not expose such content to the developer. The exact behavior depends on the model and API.
For a developer building an application through an API, the question is whether reasoning content should be preserved across turns or removed. The answer depends on the use case and the model’s API requirements. In some tool-based loops, preserving the model’s prior reasoning or other provider-specific state may be important for maintaining performance. In other APIs, reasoning content should not be copied back into the conversation. Also, incorrectly reconstructing the conversation can prevent prompt-cache reuse and increase API costs. The safest approach is to follow the model’s API documentation and model card rather than assuming that one rule applies to every model.
In the experiment done in the last post, the reasoning content was removed for every model. Unfortunately, that approach does not work well for some small language models, removing prior reasoning can cause substantial degradation in task performance, especially in tool-use loops. For the LFM2.5 checkpoints used here, retaining prior reasoning during tool-based evaluation had a substantial effect on performance.
After training, there are 3 trained models. During evaluation on the 5×5 and 10×10 grids, prior reasoning was kept. Another test was performed where previous reasoning was removed.
5x5 grid performance without previous reasoning
| Model name | InvalidTool Rate % | Failure Rate (Max moves reached) % | Efficiency (if target position was reached) | Avg Tokens |
|---|---|---|---|---|
| LFM2.5-1.2B-Gemini (GRPO) | 95.0 | 0 | N/A | 5317.5 |
| LFM2.5-1.2B-ChatGPT (GRPO) | 50.0 | 50.0 | 2.9 | 6286.2 |
| LFM2.5-1.2B-Sonnet (GRPO) | 55.0 | 45.0 | 2.89 | 6950.35 |
Note: N/A is used because the model has failed in all the rounds except 1.
10x10 grid performance without previous reasoning
| Model name | InvalidTool Rate % | Failure Rate (Max moves reached) % | Efficiency (if target position was reached) | Avg Tokens |
|---|---|---|---|---|
| LFM2.5-1.2B-Gemini (GRPO) | 100 | 0 | N/A | 4655.45 |
| LFM2.5-1.2B-ChatGPT (GRPO) | 25.0 | 25.0 | 2.8 | 16747.7 |
| LFM2.5-1.2B-Sonnet (GRPO) | 0 | 90.0 | 8.0 | 22043.2 |
Removing the previous reasoning trace caused substantial performance regression for the GRPO-trained models. The Sonnet variant still performed better than the base model on some metrics, but the other variants degraded sharply. This suggests that, for these checkpoints and this tool-use setup, retaining prior reasoning was important for maintaining task performance.
Effects on benchmark results
Since this is a small model, fine-tuning it for a task may cause regression or improvement in other unrelated tasks. An evaluation of the trained models on subsets of 5 general benchmarks is done.
The base model’s model card reports results on the full benchmark sets, whereas this experiment evaluates only sampled subsets. These results should therefore be interpreted as relative measurements within this experiment rather than as official benchmark scores.
The evaluation contains 530 questions across five benchmarks:
- 125 questions: GPQA Diamond
- 125 questions: MMLU Pro
- 125 questions: GSM8K
- 125 questions: MATH-500
- 30 questions: AIME25
Performance on sampled benchmark subsets
| Model | GPQA Diamond | MMLU Pro | GSM8K | MATH-500 | AIME25 |
|---|---|---|---|---|---|
| LFM2.5-1.2B-Thinking (Base) | 28.00 | 40.80 | 96.00 | 76.00 | 30.00 |
| LFM2.5-1.2B-Gemini (GRPO) | 33.6 | 31.2 | 93.6 | 72.0 | 36.67 |
| LFM2.5-1.2B-ChatGPT (GRPO) | 33.6 | 36.0 | 93.6 | 74.40 | 36.67 |
| LFM2.5-1.2B-Sonnet (GRPO) | 24.0 | 35.2 | 93.6 | 60.80 | 46.67 |
The Gemini- and ChatGPT-trained models show similar trade-offs: both improved on the sampled GPQA Diamond and AIME25 while decreasing on the other three tasks. The Sonnet-trained model decreased on four of the five benchmarks and increased on AIME25, where it achieved 46.67%.
Conclusion
This experiment suggests that a small 1.2B reasoning model can learn a sparse-reward interactive navigation task when successful trajectories are first introduced through SFT. The initial GRPO-only approach struggled with the cold-start problem, while SFT provided the model with successful behaviors that reinforcement learning could refine.
Simplifying the reasoning traces significantly reduced token usage while preserving strong navigation performance. Multi-turn GRPO then reduced token usage further and improved navigation efficiency relative to the corresponding SFT checkpoints under this evaluation setup.
One of the more interesting findings was the effect of removing previous reasoning. The GRPO-trained checkpoints showed substantial degradation when the previous reasoning trace was removed, suggesting that the model was using information encoded in that trace to maintain state during the tool-use loop. This does not establish that reasoning tokens are universally required, but it does show that, for these checkpoints and this environment, removing them had a large effect.
There are couple of limitations. The navigation evaluation used only 20 rounds per grid size and the benchmark evaluation used sampled subsets. Larger evaluations and multiple random seeds would make the conclusions more robust but this was beyond the available budget.
Appendix
Modal platform was used for training the models. The training was tracked using wandb.ai platform.
GitHub Repo of training and eval scripts: training and eval scripts