Back to workWork / Local AI / Lightweight ML

Local AI / Lightweight ML

Choosing a model that fits the actual task

Applied research and workflow prototypes
01 / The challenge

The challenge

Model selection needed to balance useful accuracy, latency, hardware limits and operating cost.

02 / How I approached it

How I approached it

I work with local LLMs, CPU-first inference, ONNX and GGUF alongside speech tools such as Whisper, Sherpa-ONNX and Silero. Experiments also cover OCR, embeddings, vector search, multilingual processing and TTS.

System flow
  1. Define the task

  2. Choose a model

  3. Run locally

  4. Evaluate the workflow

03 / The outcome

The outcome

A collection of applied experiments informs how I choose models and connect them to useful workflows.

04 / Tools & technologies

Tools & technologies

Continue exploring

TON Tanks

Contact

Complex challenge? Let’s make it work.

Architecture, technical leadership, or a product that needs a solid foundation.

Start a conversation

A role, a product, or a technical challenge. Share a little context.

To use the form, enable JavaScript. You can also write to me directly on Telegram.

Write on Telegram