LLM Gateway Starter
An OpenAI-compatible model gateway with routing, fallback, rate limits, usage accounting, and Docker Compose deployment.
A practical architecture note on why agent follow-ups should not be treated as plain chat, and how session, run, memory, DAG versioning, and runtime events fit together.
A practical review of why static DAG planning became heavy for general-purpose tasks, and what runtime constraints ReAct needed to work safely.
A sanitized enterprise-assessment example showing what 7B fine-tuning can and cannot solve, and how to run LoRA/QLoRA with LLaMA-Factory.
A reproducible comparison of Docling, MinerU, PaddleOCR-VL-1.6, DeepSeek-OCR 2, and dots.mocr on OmniDocBench article-300 using the same test server.
An OpenAI-compatible model gateway with routing, fallback, rate limits, usage accounting, and Docker Compose deployment.
A self-hosted LLM observability collector for traces, latency, token usage, cost, feedback, and eval events.