llama.cpp
The llama.cpp project’s new llama.app experience brings simpler installation and local model serving to its fast, hardware-flexible LLM runtime. HN discusses backend performance, multi-model serving, installation security, and whether it offers advantages over Ollama.