Reorganizing the code
In the previous version of the code I allowed claude to write up the entire logging system and this caused issues with organizing the code and keeping track of the experiments and updates. I spent a week trying to detangle that code without using Claude as per my AI policy for this project and ultimately I decided that the code had to be re-written from scratch by hand so I was able to pin-point every issue and know every part of the repo correctly. To that end I reorganized that in part 4 of the experiments. The logging was moved to a new module in log_summary/stage_02_logger.py) and the questions were turned into a class based system to enable better logging. Finally I added visualization module to visualize the metrics from multiple runs. This took longer than I expected but I think ultimately it will make further iteration easier and enable this project to be extended or be used as sub-systems for other projects.
The new questions module
Questions are now based on pydantic models BaseModel class. There is Question class which has a question type, question text and question answer. Questions are grouped into a QuestionSeries by type of questions. Eg questions that are only testing scalar arithmetic, questions which test list operations on text, questions that test list operations, questions that test combinations of operations and later expanded by more types of questions. This allows the recordings and metrics to be grouped to say where tool calls are failing and how. It also allows deeper dives into specific parts by question series.
The logging module
The logging module is now composed of two classes ResponseStatistics and ResultsLogger. ResponseStatistics is also a pydantic BaseModel class it contains the question, the series the question belongs, the answer provided by the LLM, number of total tool calls, number of invalid tool calls, number of messages for the session(based on the thread_id for the conversation), the length of the thought when reasoning is enabled, length of the response which is the length of the final message only, a boolean indicating weather the answer had the expected response, the names and arguments for invalid tool calls and the thought list.
For cases where reasoning is enabled, the length of the thought and the thought list can reveal the cases where the model tries to "deceive" the user by making a few tool calls but doing the actual calculations internally within the thought.
The main work in this module is done by the ResultsLogger class. it takes the stage of the model and the current date and time and uses those to create a folder to save the results in. The get_path_name function returns the path at which the full conversation is saved and get_statistics function builds the values required by ResponseStatistics.
A big change from the Claude generated code to this version of logging was how thought was recorded. Even using Fable5, Claude tried to get the thought values using the raw values specified by the gemma4 model card. This would work correctly if the model was being used with the Transformers library but with the LangGraph library, this is done in the AIMessage class already. When any message from MessagesState's messages arg are of type AIMessage, it can optionally contain a key called "reasoning_content" in its additional_kwargs value and this is what contains any thoughts when reasoning is turned on. Recording these values to a list indicates how speciifcally the deliberately malformed questions that needed a while loop to solve via tool alone were actually solved in reasoning and not by tool call alone.
Model changes and Graph Generator class
In this version I focused on the gemma:e4b and gemma:12b models since the 26b does not effectively fit on my 16 GB GPU. It can run but it takes a lot of time and when given the correct context size that size increases a lot more and thus each question takes a lot longer to evaluate effectively. Similarly this version fixes the parameters for top_k to 64 and top_p to 0.95 (the values recommended by the model card for gemma) by default. This can later be extended when I add meta's new glimmer and other models to the experiments.
All these changes are now encapsulated in the GraphGenerator class.
Experiment changes
The new experiments work by rrunning a grid multiple times and uses the pydantic model for summary information to keep track of the results. In this case the grid is gemma4:12b and gemma4:e4b with thinking on and off and for checking consistency it runs all questions 5 times by default.
One important detail in this script is how memory and the related state is handled. It is initialized here in the experiments and not in the graph generator class. This way each set of related questions gets one state so if one answer is supposed to inform the next calculation it can but it won't result in either a very large state cluttering the context window or in information leaking from questions and their responses to the next question when said questions were separate.
In terms of saving the results, each run produces its own statistics summary file and they are combined after all the runs in the combine_results function.