> your AI agent picks dependencies from memory; give it dated facts — try starlog.dev ↗ vet your agent's deps ↗ vibe-coding is fine. vibe-importing isn’t. — try starlog.dev ↗ vibe-importing isn’t fine ↗ your agent has never seen your private packages — try starlog.dev ↗ facts for private packages ↗ a linter for the dependencies your AI agent picks — try starlog.dev ↗ a linter for agent deps ↗ whois is redacted, cdns mask the rest — get the real operator — try whoisgeni.us ↗ who really runs that domain ↗ domain attribution that shows its work — full evidence chain — try whoisgeni.us ↗ domain intel w/ evidence ↗

← All articles

LLM Engineering

How LLM applications get built and shipped: inference and serving, structured output, token formats, evaluation, and the libraries powering production language-model systems.

157 articles

LLM Engineering

HyperQwen: 1,000 tok/s on a Single RTX 3090 Through Aggressive Recurrent-State Speculation

★ 1.4k Python Sep 17, 2026
LLM Engineering

oMLX: Why This Mac Inference Server Caches to Your SSD (And Why That's Brilliant)

★ 21.7k Python Sep 14, 2026
LLM Engineering

Qwen3.8-27B on a Single RTX 3090: How Calibrated Quantization and Speculative Decoding Hit 1,000 tok/s

★ 1.2k Python Sep 8, 2026
LLM Engineering

METATRON: Building a Penetration Testing Assistant with Local LLMs and Zero Cloud Dependencies

★ 3.6k Python Aug 24, 2026
LLM Engineering

Switchyard: Composable LLM Routing Without the Gateway Lock-In

★ 1.2k Rust Aug 13, 2026
LLM Engineering

Running a 2.78-Trillion-Parameter Model on a Laptop: Inside WASTE's NVMe-Streaming Architecture

★ 2.0k C Aug 10, 2026
LLM Engineering

Running 35B MoE Models on 8GB GPUs: How vvllm-ms Caches Experts in System RAM

★ 1 Python Aug 8, 2026
LLM Engineering

Terminal-Bench: Why Evaluating LLM Agents on Real Terminal Tasks Is Harder Than You Think

★ 2.5k Python Aug 6, 2026
LLM Engineering

OSWorld-V2: The GUI Agent Benchmark That Hides Its Answers

★ 229 Python Aug 5, 2026
LLM Engineering

Running a 2.78-Trillion-Parameter Model on a Laptop: Inside WASTE's NVMe-Streaming Engine

★ 145 C Jul 31, 2026
LLM Engineering

Running 753B MoE Models on Consumer GPUs: Hand-Written SASS Kernels for 2-Bit Experts

★ 253 Sass Jul 10, 2026
LLM Engineering

LLM Checker: Hardware-Aware Model Selection for Local Inference

★ 2.8k JavaScript Jun 29, 2026
LLM Engineering

Fighting LLM Hallucinations in 2026: llm-council vs claude-octopus vs LLM-Check

Various Jun 22, 2026
LLM Engineering

Running Large LLMs on Limited Hardware in 2026: AirLLM vs llm-checker

Various Jun 22, 2026
LLM Engineering

The Fragile Economics of Free AI: Mapping the Great Inference Subsidy Wars

★ 696 Jun 20, 2026
LLM Engineering

Shard: Proving LLM Inference Can Work Across Scattered GPUs and Terrible Internet

★ 19 Python Jun 18, 2026
LLM Engineering

Nanocoder: The Terminal Coding Agent That Lets You Switch Models Mid-Conversation

★ 2.1k TypeScript Jun 14, 2026
LLM Engineering

ds4: The SSD-Streaming Inference Engine That Treats Your Mac's NVMe Like RAM

★ 13.4k C Jun 11, 2026
LLM Engineering

Harness-1: Training Search Agents with State Externalization

★ 390 Python Jun 9, 2026
LLM Engineering

SichGate Methodology: When Healthcare CISOs Need to Red-Team 4-Bit Llama Without Hiring Offensive Security

★ 2 Python Jun 8, 2026
LLM Engineering

ModelRegression: Building a Daily LLM Benchmark That Tests What Developers Actually Use

★ 12 Python Jun 6, 2026
LLM Engineering

Inside AI Product Bench: Why Two LLMs Disagree on Half Their Product Recommendations

★ 23 HTML May 27, 2026
LLM Engineering

Neuromod-LLM: Treating Language Models Like Brains on Drugs

★ 6 Python May 24, 2026
LLM Engineering

makemore: Understanding Language Models by Implementing Them Seven Different Ways

★ 4.0k Python May 24, 2026