Skip to main content
Created By: Solomon Githu Public Project Links: https://studio.edgeimpulse.com/public/1105603/latest GitHub Repo: https://github.com/SolomonGithu/edge-ai-behavioral-coaching-assistant

Project description

Generative AI (Gen AI) models such as Large Language Models (LLMs) have become important because of their ability to generate insights and personalized feedback. AI-based coaching is becoming more recognized for improving well being across various domains. At the same time, technology advancements are making us reimage how coaching can be delivered. For example, Edge Gen AI is a topic that is creating more interest for AI researchers because of the ability to get Generative AI capabilities being executed locally on edge devices as compared to costly cloud services. However, not all edge applications require the huge computation, size and processing power like LLMs. For this reason, engineers developed Small Language Models (SLMs) to make Gen AI models more accessible by reducing the models size and computational requirements while retaining useful language understanding and generation capabilities. For comparison, GPT-3 has 175 billion parameters and requires 350GB of storage space for the weights, while Qwen 3.5 used in this project has 0.8 billion parameters, requires 507MB of storage; and you can chat with it! This project demonstrates how we can combine modern hardware and optimized software such as TinyML models and Small Language Models to create sustainable AI powered behavioral coaching assistants that run on cost effective hardware such as an Arduino UNO Q. It uses a perception layer (lightweight object detection model) to observe activities and the detections are stored over time into a simple observations list. Afterwards, the system periodically summarizes observations into a simple activity-descriptive text that is provided to an SLM for generating a short behavioral recommendation such as “Drink more water and use coffee less frequently”. Summarizing the observations enables us to reduce the amount of information passed to the SLM, reducing the prompt size and computation time. In this case, the perception and language models serve different roles: Perception layer (What is happening?) → Behavioral memory layer (Summarize what has been happening over time) → Small Language Model (Provide wellness-based recommendation)
For the current demonstration, the perception layer uses a simple object detection model that can detect a glass of water and a coffee mug in an image.
[!NOTE] I focused on object detection and temporal aggregation of observations counts rather than consumption tracking. For example, the application does not determine if a glass of water was picked, consumed and returned with less water. This is the same case for coffee mugs: detecting a mug does not mean that coffee was consumed. The perception model simply detects visible objects and records the observation. In future, there is need to advance the activity recognition to distinguish the presence of an object and interactions with it.

Components and hardware configuration

  • Arduino® UNO Q: either the 2GB or 4GB variant.
  • USB camera
  • USB-C® hub adapter with external power
  • A power supply (5 V, 3 A) for the USB hub
  • Personal computer with internet access
Software
  • Edge Impulse Studio
  • Arduino App Lab
  • Local Small Language Models in App Lab (Qwen, LLama, Gemma)
Qwen 3.5 0.8B was used for this project.

Step 1: Setup your UNO Q

First connect a USB-C Hub to the UNO Q. Next, connect a USB webcam to the Hub and power the system through the Power Delivery slot. Before working with the UNO Q for the first time, we need to setup the Linux system through the App Lab.
Arduino has documented the necessary steps for setting up the board in their user manual.

Step 2: Train an object detection model with Edge Impulse

For the perception layer, I used the Edge Impulse platform to train and deploy a lightweight object detection model to the UNO Q. Depending on the use case, you can use a model for image classification, sound classification, motion classification, anomaly detection, etc.; and reconfigure the app to process the corresponding input. We need to first collect data (images, audio recordings, motion, etc.) for training the model. I collected 100 images of a glass of water and a coffee mug on my desk. The dataset was split into 78 images for model training and the remaining 22 for testing.
The next step is to configure an Impulse which is a configuration that defines the input data type, data pre-processing algorithm, and the machine learning model training. I configured the Impulse to take images of 96x96 pixels, preprocess and normalize image data, and train a FOMO object detection model. FOMO is a lightweight object detection model based on MobileNetV2 and it is designed for constrained edge devices.
Looking at the Feature explorer, there was a good separation between the features for the two labels (coffee_mug and water_glass), indicating that the model could learn to detect the objects through transfer learning. It is also notable to mention that on the features page, the Studio highlights the estimated on-device performance for the feature generation process on the UNO Q during inference. 1 millisecond for processing an image with peak RAM consumption of 4KB is very good.
Finally, after training the FOMO model, the F1-score was 100%, which is very good. My scene was simple, consisting of a glass or mug, a keyboard, and my typing hands occasionally. This simple controlled environment contributed to the high performance of the model. The estimated on-device performance on the UNO Q was 8 milliseconds for inferencing with peak RAM usage at 133KB and flash usage of 81.3KB.
Testing the model on unseen data showed good results with 95% accuracy. This was sufficient for me to proceed with deploying the model to an Arduino UNO Q.
You can access my project here: Coffee or Water - UNO Q. Note that differences in the inferencing environment contributes to incorrect predictions from the model because it was trained with a relatively simple and few data from a specific scene. The model may not detect a glass of water or coffee mug that accurately when you test it. Differences in lighting, camera position, background, etc., can affect the features passed to the model leading to bad performance.

Step 3: Deploy detection model

Deploying the object detection model to an Arduino UNO Q is relatively simple. SSH into your UNO Q and clone this GitHub repository into the /home/arduino/ArduinoApps/ directory. This repository contains an Arduino App Lab project that implements the AI behavioral coaching assistant.
Once the repo has been cloned, you should see the project in the ‘Apps’ section.
Open your new app and click the ‘Video Object Detection’ brick. Navigate to the ‘AI models’ tab and you will see all models from your Edge Impulse projects configured for the UNO Q board. Identify your model using the Edge Impulse project name and click ‘Download’ to install it on your UNO Q. In case you face problems, Edge Impulse has documentation and a video tutorial for this step.

Step 4: Install a Small Language Model

Remember we had two cascaded models in our pipeline: a perception model (object detection for my case), and a Small Language Model. Technically, developing the first model takes the most time because SLMs are already trained for general language based tasks. Still in the new App Lab project, click the ‘Large Language Model (LLM)’ brick and navigate to the ‘AI models’ tab. Download and select one of the available models. For this project, I used Qwen 3.5 0.8B because it is relatively small and can generate a response more quickly. Other supported models such as Gemma 3 1B can also be used depending on your hardware and use case.
Using a Small Language Model enables us to have language-based reasoning running locally on the same edge device. One advantage is privacy whereby we do not need to send to the cloud sensitive information about detected objects and instructions for language processing. Secondly, local inference reduces dependency on the internet, cloud APIs, and subscription services. A small model also requires fewer computational and storage resources making it suitable for constrained edge hardware.

Application design

The application is programmed in Python with a main.py script managing the main processes: capturing frames from a camera, object detection, behavioral monitoring, prompting the SLM, and displaying the results for each process on a Web UI. A simple behavioral memory list is maintained using the label of object detections. Each detected object is stored as a dictionary containing the predicted label and timestamp. For example:
Before sending the observations to the SLM, the behavioral events are summarized. For example, individual observed labels without the timestamp, are represented as:
The application will take this kind of information and aggregates observations based on the count of each label into a simple text like:
This aggregation reduces the prompt size, requires fewer input tokens and removes unnecessary information such as timestamps and repetitions in the detection events. Prompt design is another important section in the application. The SLM is given a specific task: interpret summarized behavioral events and provide a wellness-based response.
To constrain the model from generating unnecessary text, repeat parts of the prompt, or give response which is away from the requested task, this prompt design guides the SLM on the content and format of the response.

Step 5: Run the AI coach

Start the application by clicking the ‘Run’ button on App Lab. Running the application for the first time will take some seconds since the system needs to download the necessary Docker images. Once this is finished the application’s container will be started and the Web UI will automatically open in a browser. You can also open the Web UI manually in a browser by setting URL to the local IP address of your Arduino UNO Q and port 7000.
Periodically place a glass of water or coffee mug in the camera’s field of view. After sometime (defined with the AI_COACHING_INTERVAL_SECONDS variable in main.py) the SLM generates a recommendation based on the observations since the previous coaching cycle. Below is a snapshot of a response from the AI coach when a glass of water was detected more frequently. Though small, the language model has both an understanding of the relationship between water and health, plus vocabulary such as ‘hydration routine’. The response time from Qwen 3.5 was around 60 seconds, with peak CPU utilization approximately 95% while RAM consumption was 1.25GB out of the available 3.58GB.
In another test, a coffee mug was mostly detected, ‘simulating’ more frequent coffee consumption. Qwen 3.5 0.8B was able to relate that coffee and caffeine are a topic of discussion. The coach suggested to limit coffee intake so as to maintain focus and prevent dependency.

Too much coffee

Coffee or water? This time the perception layer observed approximately equal observations of both a coffee mug and glass of water. The analysis from the AI coach was good, that ‘Frequent hydration vs occasional mugs indicates a strong focus on fluid intake’. Recommendation was also relevant: ‘Prioritize water consumption to ensure full hydration’. This demonstrates an important characteristic of generative AI whereby the output from the models is dynamic based on the contextual instruction given to the model as compared to having a limited set of predefined responses.

Coffee or water

Because of the general limitation of the model, incorrect responses can be generated. For example, in this test the perception layer observed more coffee mugs than glasses of water. The analysis from the model was correct, that ‘The user prefers coffee over water’. However, the recommendation was misleading: ‘Replace water glasses with coffee mugs for 1 week’. The response was grammatically correct but not valid from a wellness objective. This limitation comes from several factors such as the ‘small’ size of the model, ambiguity in the instructions and mostly the generative nature of language models.

Bad recommendation

Conclusion

This project has demonstrated creating Edge AI coaches using a cascading architecture of having a lightweight perception model and a Small Language Model. The system divides the task into stages to reduce computational requirements. The approach also shows how easy it has become to develop and deploy small perception models using a platform like Edge Impulse and leveraging open-weight (free) language models such as Qwen 3.5 for language-based reasoning. However, there are several future advancements that can be made to the project. The first one is to implement a robust approach to replace observation counting with interaction awareness. Instead of treating an event such as coffee detection is treated as one observation, the system can combine multiple observations to identify a more meaningful interaction, that is, coffee mug appears in frame → person’s hand on mug → mug disappears → mug appears with less coffee; is treated as one observation. Finally, the system can also learn a personal habit of each user such as “This person drinks water frequently…”. This can assist in identifying deviations from the known habit and provide a better context-aware recommendation.