Host the model. Govern the service.
Deploy an Ollama-backed AI service you control, then build a small control plane for authentication, model policy, observability, and repeatable evaluations.
A model endpoint is not a platform.
Begin local-first. Ollama listens only on loopback or a private network. An authenticated API gateway is the sole application entry point. The gateway enforces identity, rate limits, model allowlists, request limits, and safe telemetry before forwarding to Ollama.
Serving plane
Client → HTTPS gateway → authorization and policy → Ollama model runner. Store no prompts by default; expose health and latency metrics without raw content. Pin model names and record their provenance.
Control plane
Admin-only model catalog, per-role quotas, policy version, deployment status, and evaluation results. Separate admin and learner views. Every policy change has an auditable owner and rollback.
From local demo to defensible service.
Run privately
Install Ollama from its official source and pull a model that fits your hardware and license constraints. Bind locally first.
Front it
Build an API service with authenticated sessions, authorization, size and rate limits, and request IDs.
Operate it
Add model allowlisting, bounded concurrency, health checks, sanitized logs, and an admin policy view.
Evaluate it
Run a versioned suite for normal tasks, refusal boundaries, outages, and unauthorized model access.
What must work before you call it hosted.
GET /api/models role-filtered catalog
GET /api/health service health without private details
GET /api/metrics sanitized operational counters
POST /api/evals admin-only, versioned test suite
- Ollama endpoint not directly internet-exposed
- Authentication and role tests pass
- Rate and concurrency limits demonstrated
- Model/license/provenance documented
- Prompt logging off or redacted by policy
- Failure, timeout, and rollback test captured
Portfolio package
Provide an architecture diagram, reproducible local deployment instructions, API contract, threat model, safe demo account, evaluation results, and a short operations runbook. Redact tokens and user content.
Interview prompt
Why put a gateway in front of Ollama? How do model choice, latency, privacy, and abuse controls affect your design? What breaks when the model runner is unavailable?
Defend the model and the controls around it.
Present what you built, what failed, how you verified it, and what you would improve.

