Sohaib Arif — AI Safety Research

Notes on AI safety for local & embedded models

About this blog

I am a software research engineer with 8+ years of experience. I have worked with cutting edge technologies and helped research and apply them, built prototypes and proof of concept projects, and taken them to production. I love researching topics that are new and ambiguous. Most of my experience is related to machine learning and deep learning, but I have worked with several other related and unrelated topics going from explainable and interpretable AI/ML, to making quantized ML models on embedded systems to a brief dive on Quantum Computing and researching its impacts on software security. Most recently I have been working on AI safety with a particular emphasis on local generative AI. I am using this blog to track some of my currently running side projects and learnings starting with the LLM tool use verification project.

Projects Tracked in this blog

Local LLM Tool Use Verification

What happens when you ask an LLM to do some calculation? Most people would assume it would just write out the formula and the run it and then write the result. That is not the case. The most famous example of non-technical people discovering this was when in August 2024 researchers found out that LLMs can't tell the number of times the character 'r' appears in the word strawberry. The reason was that LLMs were not doing any calculation internally to provide the answer, they had tokenized the word into a number in some way without splitting it into its component characters and when asked about a calculation having to do with its spelling, it just used the whole word and output what it "thought" was right. In that case it was mostly inconsequential but as more tasks get delegated to LLMs including local LLMs running on embedded devices or SOCs these kinds of calculations need to be done in code only instead of in memory. This project will track if local LLMs use provided code to answer questions correct and what failure modes they produce when they don't. Some modes I have seen failure or arguable failure is the LLM using its own memory to answer, using a combination of correct tool use and its memory to correctly answer, and some cases where it called a tool but didn't use its answer. This project will track how and when failures occur to use tools and what kind of behaviour can be observed with different types of models and parameters.

Multi-Agent communication

This project will track safety issues with multi agent communication. Most AI safety so far has focused on single agents but systems can now be built with multi-agent communication in mind. Sometimes they are supposed to communicate but other times they are supposed to work independently. Both are LLM safety issues. In cases where they are supposed to communicate, problems are related to miscommunication, underspecified or overspecified goals, or conflicting/hidden/misunderstood goals. Then there are cases where agents are not supposed to communicate such as when one agent is being used to train or evaluate another. In those cases research has found communication being conducted via hidden channels. This project builds on the previous one to explore the these issues in local LLM models.

Workout Tracking using local LLM

This project will expand on my submission for the Kaggle Gemma4Good competition submission. In that submission I created a project to analyse workouts using only local LLMS keeping the data private and inference costs low so it could be used in physical therapy or high end sports training or any other cases where privacy needs to be preserved and where the user does not want to pay the monthly and usage fees for commercial models. What I discovered early in this project was that the analysis can be wrong simply because the question was not asked correctly in a chatbot or the model used the wrong formula. So I pivoted the original project for robustness by grounding the answers in only specified books and research papers with the main meat of the project being a custom 3 step RAG pipeline involving reading the papers, reading specific formulas from the papers and then generating code based on those formulas. The remaining project could not be done in time for the competition submission both because the custom RAG pipeline took a long time to complete and debug and also because the resulting function calls needed to be verified correctly which is the mentioned LLM Tool Use Verification project being tracked in this blog. This project will resume work on the actual workout tracking and more sensors and depth estimation using Intel Realsense in combination with the IMU sensor on the Arduino BLE Sense 33 data.

Questions or comments are welcome at arif_sohaib@outlook.com.