The llamay blog

Long-form on running language models on hardware you own: local inference, prompt caching, GGUF quantisation, tool calling, structured output, and what any of it actually costs once the tokens are yours.

RSS · one article a day, written on llamay.