ProjectsA.L.E.K.S.Y v1

A.L.E.K.S.Y v1

A fully offline Polish voice assistant. No cloud, no API keys, nothing leaves the device.

A Polish voice assistant where every stage runs on the device: wake word, speech recognition, the language model and speech synthesis. It runs on an NVIDIA Jetson Xavier NX inside an orange 3D-printed case with built-in speakers, powered from a USB-C power bank.

Backstory

A.L.E.K.S.Y started in a computer security course. The assignment was to design a secure system on paper, but I thought it was meant to be real, so we started building one. The idea was a voice assistant with a strong privacy guarantee: if nothing is ever sent anywhere, there is nothing to intercept. A Raspberry Pi 5 with an AI HAT was not enough, so we moved to an NVIDIA Jetson Xavier NX borrowed from the university. The name came from our friend Aleksy, because Amazon Alexa sounded like him. A teammate set up the first wake word code, and from there the build was mostly mine. The Jetson quickly turned out to be the weakest part of the whole build. It worked, we got top marks, and I happily gave the Jetson back. What went wrong is below, and it is the reason v2 moved the heavy work to a server.

Problems along the way
  • The Jetson Xavier NX is bad at LLMs. On a 7B model quantized to 3 bits it made 1 to 2 tokens per second, so a single answer took 20 to 30 seconds. My MacBook Air beat it at everything.
  • JetPack 5 locks the board to an old software stack. Ollama was stuck at 0.1.46 without tool calling, onnxruntime had to be pinned, and newer Polish models like Bielik v3 did not run at all. The best working option was Bielik 7B v0.1 in Q3_K_M, which could barely hold a conversation or keep track of much context.
  • Speech-to-text was just as slow. The GPU was reserved for the LLM, so Whisper ran on the CPU and needed 5 to 8 seconds to transcribe a 2-second sentence.
  • The Raspberry Pi 5 with an AI HAT, the first plan, could not run an LLM at all. The HAT accelerates vision models, not language models.
  • The audio was improvised. The lavalier microphone ran on batteries and was flat the morning after I left it on, and the speakers came from a seven-year-old project with an amplifier that hissed whenever the board worked hard.
Takeaways
  • Check what software a board actually supports before you buy into it. Specs on paper meant nothing when the newest Ollama and models would not install.
  • An edge board from 2019 is not an LLM machine. For a conversation that feels natural you need far faster inference than it can give.
  • A fully offline device is a real privacy guarantee, but it caps the quality at what the local hardware can run. That trade-off is why v2 moved the heavy work to a server.
  • Latency is the whole experience. A 20-second wait kills a voice assistant no matter how good the answer is.
  • Audio hardware is not an afterthought. A cheap microphone and a noisy amplifier make even good answers sound bad.
Highlights

Fully Offline

Wake word, STT, LLM and TTS all run on the Jetson

Polish End to End

Bielik LLM through Ollama, Piper TTS with a Polish voice

Conversation Memory

Last five exchanges, kept in RAM only

Voice Timers & Self-Awareness

Sets timers and reports its own CPU temperature and RAM

Edge Hardware

Jetson Xavier NX in a custom case, on a power bank

Demo

Screenshots