Visualizing Results
After the previous updates I needed ways to visualize the outputs and keeping with my AI policy for this project I didn't use LLMs except for ideation. First I had to decide what to visualize. First thing was correctness i.e. which models and in which mode gave the most correct answers. Correctness in this case was measured by the logging module checking weather the final message output for each question contained the answer.
In this case the parameters were fixed except for turning thinking on or off. There were 2 models (gemma:e4b and gemma:12b) with 2 modes (reasoning on and reasoning off). The thoughts were collected by the logging module and the entire conversations were saved in folders marked by the date/time when the experiment was run and the settings used. To confirm that answers were the same or similar across runs, each experiment was run 5 times for each of the 4 settings. There were 5 total questions, with the complex question being 2 questions in one group so in total there were 2 x 2 x 5 x 4 = 80 files of conversations, 4 x 5 files of summary statistics and 5 files of joined joined results to visualize.
Grouped model name and reasoning to check if each group contained the answer. For this visualization I grouped files by model name and reasoning and summed the number of times the final message had the expected answer. To get the joined results for the last n runs I used rglob (recursive version of glob) and then reverse sorted by the stem of the parent directory. The parent directory in this case was the one named after the date and time of the run so reverse sorting it by its name gives the last n directories.
joined_stats_across_runs = list(Path("results","stage_04").rglob("*joined_results.csv"))
joined_paths_last_5 = sorted(joined_stats_across_runs,key=lambda x:x.parent.stem)[-5:]
Then I need to group the results across these runs by question, model_name and reasoning but since the grouping is across 5 files, pandas's groupby function can't be used directly. Instead I first read in all files a dictionary of data frames
summary_data = {path.parent.stem: pd.read_csv(path) for path in data_paths}
Then concatnate these dataframes by using the date time extracted from the folder name as a new column
df = pd.concat(
[d.assign(run=name) for name, d in summary_data.items()],
ignore_index=True,
)
groupby still can't be used because the columns are still mixed types some strings, ints and booleans.
invalid_tool_names and invalid_tool_args are lists and internally in the csv they appear as strings of lists so python's abstract syntax tree (ast) library's eval function function can be used to turn them from strings into lists of strings in case they were not being read as lists already.
for col in ["invalid_tool_names", "invalid_tool_args"]:
df[col] = df[col].apply(
lambda v: v if isinstance(v, list)
else ast.literal_eval(v) if isinstance(v, str) and v.startswith("[")
else []
)
The answer_had_response needs to be converted into a boolean so the true values can be counted
df["answer_had_response"] = df["answer_had_response"].astype(bool)
Finally a custom aggregation can be used on all columns.
keys = ["model_name", "reasoning", "question_series", "question"]
int_cols = ["num_tool_calls", "num_invalid_calls", "msg_count", "len_thought", "len_response"]
agg = df.groupby(keys, as_index=False).agg(
**{c: (c, "sum") for c in int_cols},
n_runs=("run", "nunique"),
n_correct=("answer_had_response", "sum"),
accuracy=("answer_had_response", "mean"),
invalid_tool_names=("invalid_tool_names", lambda s: [x for lst in s for x in lst]),
invalid_tool_args=("invalid_tool_args", lambda s: [x for lst in s for x in lst]),
)
Now these aggregated summaries can be visualized as barplots using seaborn (dataframe also has its own version of barplot but for this use case searborn's hue parameter is needed).
agg_sum = pd.DataFrame(agg_data.groupby(['model_name','reasoning'])['n_correct'].sum()).reset_index()
agg_sum['model_name_full'] = agg_sum['model_name'] + agg_sum['reasoning']
sns.barplot(agg_sum, x='model_name', y='n_correct', hue="reasoning")
Across 5 runs here are the results:

The surprising thing is that gemma4:e4b both reasoning and non-reasoning gets more accurate answers at least by the criteria measured. I say "by the criteria measured" because in this case I am not counting fuzzy matching as correct answers and I will add that in a later update.
We can also zoom in and compare two specific runs with a simple count based match and grouping.
df = pd.read_csv(data_path)
grouped_df:pd.DataFrame = df.groupby(['model_name','reasoning'])['answer_had_response'].sum()
Specifically the last two runs return the following counts.

you can see in these results that the two runs did not produce the same counts but for both cases gemma4:e4b with reasoning got the most correct answers.
With the grouped results we can also check what impact the length of reasoning and number of tool calls had on accuracy. We can compare this against the length of the response overall vs length of reasoning to check if longer response + long reasoning or short response + long reasoning lead to more accuracy. In the graph on the left, the one on the left represents has the x axis representing the number of tool calls overall, the y axis represents the length of thoughts and the hue is the name of the model, note that in this case because the length of thoughts is being measured, the metric is only for runs where reasoning was turned on. The graph on the right shows the length of thoughts compared to length of response measuring weather in cases where thinking is turned off, does the model just add that to the response to get similar results.

The next visualization measures weather models could recover from incorrect tool calls to still provide the correct answer by just adding more messages eg. one message sends bad tool call name or args, next reads the error and provides correct tool name or args and then the result recovers. Notice that the value being measured here is the number of messages, not the length of the response.

From these results we see that there were no invalid tool calls but that very high number of messages (which include tool calls) does not necessarily increase accuracy. gemma4:e4b reasoning and non reasoning show the highest number of correct responses. It can use up to around 25 messages to get all 5 answers correct while gemma4 12b used up to 65 messages to get 3 correct answers. Also remember that this graph is the distribution of all runs so the numbers above are the more values on the edges. However it should be noted that for the complex question, using less messages could be an indicator that the model did respond accurately but did the calculation in reasoning alone and not with the tool. This is because those two questions required using the tool several times to create a custom loop.
Changing the Y axis to number of tool calls and hue to the number of correct answers we get the following graph

This graph is a better indicator of what I said earlier. 12b reasoning has the highest number of messages and most of those are tool calls, but it is getting 3 correct instead of 5. Similarly for e4b, we get 5 correct answers and 25 messages, most of which are tool call messages. You can also see 5 correct with less than 10 messages which may need further examination to check if the solution was done in memory vs via tool call.
The last 2 visualizations are word counts in all conversations that were recorded. First we get all of the conversation files. These are in the results folder and saved as txt with parent folders indicating weather reasoning was on or off and what the temperature and context size were (both of which were kept static for these experiments).
def get_dirs_for_full_conv(stage:str, last_n:int)->list[Path]:
"""
return the directories containing the last_n runs for a provided experiment stage
"""
base_path = Path(Path.cwd(),"results", stage)
if Path.cwd().stem == "visualize_results":
base_path = Path(Path.cwd().parent,"results", stage)
dirs_in_base = [curr_dir for curr_dir in base_path.iterdir()]
dirs_in_base = sorted(dirs_in_base, reverse=True)[:last_n]
return dirs_in_base
For the word cloud visualization, I am using the wordcloud module with all the known stopwords and some added stopwords.
dirs_in_base = get_dirs_for_full_conv(stage, last_n)
print([dir_path.stem for dir_path in dirs_in_base])
conversation_dict:dict[str, list[str]] = defaultdict(list)
custom_stopwords = set(STOPWORDS)
custom_stopwords.update(["ai","human","tool","tools"])
Then going over each text file and saving it to a dictionary with the name of the question as key:
for folder_name in dirs_in_base:
conversations = list(folder_name.rglob("*.txt"))
print(folder_name.stem)
for conversation in conversations:
with open(conversation,'r') as f:
conv_text = f.read()
conversation_dict[conversation.stem].append(conv_text)
And then finally the word counts for each conversation can be calculated.
for key, val in conversation_dict.items():
print(key)
val_merged = "\n".join(val)
wordcloud = WordCloud( # pyright: ignore[reportUnknownMemberType]
background_color='black',
stopwords=custom_stopwords,
colormap="viridis").generate(val_merged)
plt.imshow(wordcloud, interpolation='bilinear') # pyright: ignore[reportUnknownMemberType]
plt.title(key) # pyright: ignore[reportUnknownMemberType]
plt.axis("off")# pyright: ignore[reportUnknownMemberType]
plt.imsave(Path("images","stage_04_visualizations",f"wordcount_for_{key}.png"), wordcloud)# pyright: ignore[reportUnknownMemberType]
Here are the aggregated word clouds for all 5 questions:
complex multi step list question

multi step scalar question

scalar arithmetic question

string list question

The last visualization for this blog post is the difference in word counts between reasoning and reasoning off. For this version the basic wordcount library does not work so instead python's built in Counter class from collections module needs to be used
This following function gets all the words for each response
def get_all_responses_for_question(question:str, response_paths:list[Path]):
"""
get all responses for question
"""
all_response_versions:list[Path] = []
for path in response_paths:
all_response_versions += path.rglob(f"*{question}.txt")
word_frequencies:dict[str, Counter[str]] = {}
ignore_words:list[str] = ["*"*10,"ai","human","tool","the","The"]
for full_file_path in all_response_versions:
version_name = "_".join(str(full_file_path).split("/")[-6:-1])
with open(full_file_path,'r') as f:
text = f.read()
word_frequencies[version_name] = get_word_frequency(text,ignore_words)
return word_frequencies
And this function does the manual word count while removing all stop words and custom words mentioned in the previous function
def get_word_frequency(text:str, ignore_words:list[str])->Counter[str]:
lines:list[str] = text.split("\n")
all_words:list[str] = []
for line in lines:
all_words += line.split()
custom_stopwords = set(STOPWORDS)
custom_stopwords.update(ignore_words)
word_count:Counter[str] = Counter(all_words)
for key in custom_stopwords:
if key in word_count:
word_count.pop(key)
return word_count
And finally this one visualizes the differences between the average word counts for a particular questin's response messages compared to a specific run based on the model name and reasoning being on or off.
def visualize_wordcloud_diff(stage:str, last_n:int, question:str,model_name:str, thinking:bool):
"""
visualize wordclouds for differences in the same conversation across different runs
"""
dirs_in_base = get_dirs_for_full_conv(stage, last_n)
print([dir_path.stem for dir_path in dirs_in_base])
word_frequencies = get_all_responses_for_question(question,dirs_in_base)
total_counter:Counter[str] = Counter()
for counter in word_frequencies.values():
total_counter += counter
n = len(word_frequencies)
avg_counter_counts:dict[str, float] ={word: count / n for word, count in total_counter.items()}
avg_counter:Counter[str] = Counter(avg_counter_counts)
sample_keys = list(word_frequencies.keys())
if thinking == False:
filter_samples = list(filter(lambda x: model_name in x and "think_off" in x, sample_keys))
else:
filter_samples = list(filter(lambda x: model_name in x and "think_on" in x, sample_keys))
print(f"{filter_samples=}")
sample = random.sample(filter_samples,1)[0]
sample_diff = word_frequencies[sample] - avg_counter
_, (ax0, ax1) = plt.subplots(nrows=2, ncols=1, figsize=(10,8)) # pyright: ignore[reportAny]
sample_diff_dict_df:pd.DataFrame = get_most_common_df(sample_diff,10)
ax0:axes.Axes = sns.barplot(sample_diff_dict_df, y='word',x='frequency', ax=ax0)
ax0.set_title(f"difference from avg word count for sample experiment\n {sample}") # pyright: ignore[reportUnknownMemberType]
avg_counter_df = get_most_common_df(avg_counter,10)
ax1:axes.Axes = sns.barplot(avg_counter_df, y='word',x='frequency', ax=ax1)
ax1.set_title("avg word counts across experiments") # pyright: ignore[reportUnknownMemberType]
plt.tight_layout()
plt.savefig(Path("images","stage_04_visualizations",f"wordcount_diff_{question}_{model_name}_{thinking}.png")) # pyright: ignore[reportUnknownMemberType]
Here are two results from gemma 12b with reasoning on vs reasoning off

When reasoning is off, you can see that the calculation is being done in messages based on the 9,410 and 9,440 being present in word count but you don't see these words being represented in the average word counts across experiments for the same question. This is because when reasoning was on, the same calculation was done in reasoning memory.
To verify the above claim, you can see the difference for the same question when reasoning is turned on and then the difference from average becomes less in the types of words being shown.
Conclusion
From these visualizations we can see that measuring weather or not tools were used correctly is not a straightforward task. The simple visualization at the start may indicate that e4b with and without reasoning may actually be producing the full answers most often but then a deeper dive reveals issues such as formatting difference from the original question vs 12b's answer or number of tool calls being low indicating that calculations were done in memory instead of using the required python tools. We can also see some potential recovery for thinking going on with the larger number of tool calls with 12b even if the answer is not exactly in the same format we may expect, it is still using tools to try to calculate it instead of in-memory work.