Local AI & software architecture consultancy

AI that runs on your hardware, not someone else's.

I help companies deploy private, self-hosted AI — open-weight LLMs on your own servers, on-device models inside your apps, and ML pipelines that keep sensitive data in-house. As an ex-FAANG tech lead, I also help teams design and review the software around them.

Built with
  • vLLM
  • llama.cpp
  • Ollama
  • MLX
  • Core ML
  • PyTorch
  • ONNX
  • YOLO
  • Docker
Services

Everything between "we should use AI" and a system in production.

Hands-on engineering, not just strategy. I design, build and deploy — then make sure your team can run it.

Self-hosted LLMs

Open-weight models served on your own GPUs or Apple Silicon, behind an OpenAI-compatible API your team already knows how to use.

On-device AI for apps

Quantized models running directly on iPhone, iPad, Mac and Android — summarization, chat and vision with no server round-trip.

Private RAG & search

Question-answering and semantic search over internal documents, mail and scans — indexed, embedded and served inside your network.

Agents on open models

Tool-using agents that run on local models, with routing that escalates to a frontier API only when a task genuinely needs it.

Computer vision pipelines

Detection, classification and tagging for images and video, from dataset labelling through training to batch or real-time inference.

Hardware & cost sizing

Which model, which quantization, which box. Benchmarks on your workload so you buy the hardware you need — and not more.

Software architecture

Systems that are still easy to change a year from now.

As a tech lead at a FAANG company, I designed and reviewed production systems and the design docs behind them. I bring the same habits to teams of any size: write the design down, make the trade-offs explicit, and plan for how it fails before it ships.

Ex-FAANG tech lead

Architecture & design reviews

A second opinion on a design doc or an existing system before it gets expensive to change: scaling limits, failure modes, data model and running cost.

System design for new products

Service boundaries, APIs, storage and deployment shaped around what you need now, with a clear path for when you grow.

Tech debt & modernization

Which parts to fix, which to leave alone and in what order, so the system improves without stopping feature work.

Fractional tech lead

Hands-on technical leadership for teams without a senior lead: design reviews, code reviews, mentoring and engineering practices that stick.

Why local

Frontier APIs are great. Sometimes they're the wrong tool.

Open-weight models now handle a large share of real business workloads. When privacy, cost or latency matter, running them yourself is often the better trade — and I'll tell you honestly when it isn't.

01

Your data stays put

Prompts, documents and outputs never leave infrastructure you control — simpler GDPR and client-confidentiality conversations.

02

Predictable cost

A fixed hardware budget instead of per-token bills that grow with every new user and every longer context.

03

Low latency, works offline

Inference next to the user or on the device itself. No network dependency, no third-party outage taking you down.

04

No lock-in

Open-weight models and standard APIs. Swap models as better ones ship, without rewriting your product.

Approach

A short path from idea to evidence.

Every engagement is built around measurable results on your own data, early.

  1. 01

    Assess

    We map the use case, data sensitivity, latency targets and budget, and decide whether local is the right call at all.

  2. 02

    Prototype

    A working proof of concept on your data, with measured quality, speed and cost — not a slide deck.

  3. 03

    Deploy

    Production serving, monitoring and evaluation set up on your hardware, cloud account or devices.

  4. 04

    Hand over

    Documentation, runbooks and a walkthrough so your team owns and operates it from day one.

Engagements

Pick the depth that fits.

Fixed-scope where possible, so you know what you're getting. Ongoing retainers available after delivery.

Advisory

A focused session to pressure-test a plan.

  • Architecture, system design or model selection review
  • Hardware and hosting recommendation
  • Written summary with next steps
Get in touch

Build & deploy

From prototype to something people rely on.

  • Production deployment
  • Evaluation & monitoring
  • Team handover and support
Get in touch
About

Hi, I'm Alexandru Tudorica.

I'm a software engineer and former tech lead at a FAANG company, where I led the design of production systems and reviewed other teams' architecture. I've spent years building systems end to end, from backend services and infrastructure to native mobile apps. These days I focus on making modern AI practical to run privately: serving open models on workstations and servers, shipping on-device inference in iOS apps, and wiring up the pipelines around them. I also take on software architecture work on its own.

You work directly with me, from the first call to the handover. No account managers, no junior hand-offs.

Have a use case in mind?

Send a few lines about what you're trying to do, what data is involved and any constraints. I'll reply within two business days, usually with a few questions and a suggested next step.