Local AI / Lightweight ML
Choosing a model that fits the actual task
Applied research and workflow prototypes
The challenge
Model selection needed to balance useful accuracy, latency, hardware limits and operating cost.
How I approached it
I work with local LLMs, CPU-first inference, ONNX and GGUF alongside speech tools such as Whisper, Sherpa-ONNX and Silero. Experiments also cover OCR, embeddings, vector search, multilingual processing and TTS.
Define the task
Choose a model
Run locally
Evaluate the workflow
The outcome
A collection of applied experiments informs how I choose models and connect them to useful workflows.
Tools & technologies
Continue exploring
TON Tanks
Contact
Complex challenge? Let’s make it work.
Start a conversation
A role, a product, or a technical challenge. Share a little context.
To use the form, enable JavaScript. You can also write to me directly on Telegram.
Write on Telegram