Links
Stack: Gemma 3 1B QAT · Ollama · FastAPI · Cloud Run
API status
Checking…
Option 1 — Intent classification
Typical latency on Cloud Run: ~8–11 s (warm instance).
—
Option 1 — Lesson plan generation
Typical latency on Cloud Run: ~3 min for CUR-NG-001 (sequential LLM calls). Leave this tab open.
—
Option 2 docs (in repo)
Offline LLM design · Multilingual roadmap · Stakeholder email · Code review
See the docs/ folder in the GitHub repository for full Option 2 deliverables and architecture charts.