Curated developer articles, tutorials, and guides — auto-updated hourly


My desktop runs two mismatched GPUs: a 20GB Ampere card and a 16GB Pascal Tesla. For months, a 21GB....


Stop trusting your local LLM's JSON. Combine Ollama schema-constrained decoding with a resilient par...


I Replaced All My Cloud AI With Local Models — Here's What Actually Broke I write a lot...


Starting work on the next Flash Onyx. Two new domains in the system prompt, and why I cut both of th...


Your local LLMs feel dumb? I fixed local LLM quality by combining a context-stacking prompt techniqu...


I posted Onyx 2 yesterday. Then I found out why it felt so slow, and the answer was sitting in the s...


Context: Building the vector store in the last entry was only half the job — the actual point of...


Context: Chunking means splitting a long document into smaller overlapping pieces before embedding.....


A post by ghost


Ollama v0.33 เปิดตัว, ใช้ Claude Desktop กับโมเดล local ได้ด้วย toggle เดียว โดย Nokka...


The Problem With Cloud AI Every token costs money. Every API call adds up. And your data...

My 5090-rig eval uses four output caps: 50, 180, 192, and 256 tokens. A score without the cap is not...

A failed local LLM row marks the test boundary. My RTX 5090 report shows why quality, speed, and set...


Lập trình viên và kỹ sư phần mềm thường xuyên phải đối soát tài liệu kỹ thuật, API nội bộ hoặc mã...

My RTX 5090 test shows how watts and output rate become joules per token, and why the faster of two ...


Nhiều lập trình viên thường tập trung vào xung nhịp vi xử lý mà bỏ quên bài toán tắc nghẽn bộ nhớ kh...

Gemma 4 26B Q4_K_M averaged 49 W on a long RTX 5090 run and peaked at 338 W. Keep both watt numbers ...


ตั้งค่า Opencodex + Ollama บน MacBook, ใช้ Claude Code และ Codex กับโมเดล local ฟรี โดย...


Production setup: Ollama + free APIs + screen sessions + daemon monitoring