Eieye
Eieye
FeedChatReportsAbout
ModelsAgentsCoding toolsOpen sourceResearchInfrastructureAI applicationsSafety and policyBusiness and industry
Filter stories by topic
ModelsAgentsCoding toolsOpen sourceResearchInfrastructureAI applicationsSafety and policyBusiness and industry

Updated 1 Sept, 17:26

14 Aug 2026
ModelsAgentsOpen sourceInfrastructure

HashAgent – Share an AI agent as a URL, runs locally via WebGPU

HashAgent packages an AI agent into a shareable URL and runs its LLM locally in the browser via WebGPU, with optional tools and offline caching. The project highlights the privacy and hosting advantages—and the practical model and memory limits—of browser-based agents.

HN Discussion
14 Aug 2026
InfrastructureAI applications

A simple fix for LLM tail latency

An LLM voice-agent developer cuts worst-case response time by sending duplicate requests and taking the faster result, rivaling a 2× priority tier. HN discusses cheaper threshold-based hedged requests, caching, and the cost and load-balancing tradeoffs.

HN Discussion
14 Aug 2026
Open sourceInfrastructureAI applications

Show HN: Lumabri – Run Moe Models on a P2P Swarm with Colibri

Lumabri is an Apache-licensed P2P inference system that spreads large mixture-of-experts models across CPUs and GPUs, transferring only needed weights or activations. HN discussion highlights its LAN use cases, latency and throughput tradeoffs, and the difficulty of guaranteeing deterministic results across hardware.

HN Discussion

Built by Will Etheridge

wjeth.comwjeth@pm.me
13 Aug 2026
AgentsCoding toolsInfrastructure

Show HN: MCP-stama – An ultra-fast Rust MCP server with no dependencies

mcp-stama is a dependency-free Rust MCP server providing fast file search, Git, and Docker tools to AI coding agents. It emphasizes sub-millisecond responses, low memory use, and simple editor configuration.

HN Discussion
13 Aug 2026
ModelsOpen sourceInfrastructureAI applications

Chestnut – eGPU dock with open-source firmware

Comma.ai’s Chestnut is an open-source-firmware USB4 eGPU dock designed to run substantially larger openpilot driving models in cars. HN discusses its tinygrad integration, hardware limitations, cost, and safety fallback behavior.

HN Discussion
13 Aug 2026
Infrastructure

Terabytes of credentials leaked in supply-chain attack

A supply-chain compromise of LiteLLM exposed credentials from roughly 2,500 organizations and 434,000 CI/CD pipelines. The breach highlights the systemic risk of rushing AI tooling into software supply chains and failing to rotate secrets.

HN Discussion
13 Aug 2026
InfrastructureSafety and policy

AI Is Threatening Natural Resources for Billions

A UN-linked report warns that AI’s rapidly expanding data-center footprint is putting pressure on water, energy, land, and minerals. HN debates the scale of those costs, comparisons with other industries, and whether stronger measurement and regulation are needed.

HN Discussion
13 Aug 2026
ModelsAgentsInfrastructureBusiness and industry

Accelerating GPT-5.6 Sol Ultrafast

OpenAI and Cerebras are previewing Ultrafast, a GPT-5.6 Sol API tier delivering up to 750 output tokens per second. HN debates its benchmark claims, likely premium pricing, Cerebras’s wafer-scale architecture, and how low-latency inference could change agentic coding and real-time applications.

HN Discussion
13 Aug 2026
ModelsInfrastructureBusiness and industry

Previewing Ultrafast mode: GPT‑5.6 Sol at up to 14X the speed

OpenAI is previewing GPT-5.6 Sol’s Ultrafast mode, promising speeds up to 14× faster through Cerebras hardware. HN discussion focuses on the infrastructure partnership and likely price premium.

HN Discussion
13 Aug 2026
ModelsOpen sourceInfrastructure

AI At Home Part 1: A Box Of Scraps

A developer builds a four-GPU home inference server from used AMD hardware, custom cooling, and salvaged parts. HN discusses ROCm support, local model performance, and whether self-hosting is worth the cost versus cloud APIs.

HN Discussion
13 Aug 2026
AgentsOpen sourceInfrastructure

We eliminated 1,400 CVEs in NanoClaw's container images

Echo describes reducing roughly 99% of the reported CVEs in NanoClaw’s container through dependency upgrades, OS package patching, and AI-assisted security backports. HN debates whether the raw CVE count is meaningful and whether agent-runtime security claims address prompt injection.

HN Discussion
13 Aug 2026
ModelsOpen sourceInfrastructureAI applications

I built a 500k-domain search engine for makers in a weekend for $10

An open-source personal search engine uses a small local language model to summarize and categorize 560,000 domains for about $10 in GPU time. The build report and discussion examine crawl steering, model-generated taxonomy problems, local inference costs, and AI-assisted content concerns.

HN Discussion
13 Aug 2026
ModelsInfrastructureBusiness and industry

Nvidia doubles RTX PRO 6000 Blackwell's MSRP to a staggering $16,000

HN discusses Nvidia’s $16,000 RTX PRO 6000 Blackwell largely as an LLM inference accelerator, comparing its memory, performance, and CUDA ecosystem with Mac Studios and multi-GPU AMD systems. The debate centers on whether its capability justifies the steep premium and how pricing may shape local AI access.

HN Discussion
12 Aug 2026
ModelsAgentsOpen sourceInfrastructure

Qwen3.8-2.4T

Qwen releases FP8-quantized weights for its 2.4T-parameter Qwen3.8 model, with 95B parameters activated and support for vLLM, SGLang, and other inference stacks. HN discusses its agent and coding benchmarks, upcoming smaller variants, and the enormous hardware required to run it.

HN Discussion
12 Aug 2026
ModelsOpen sourceResearchInfrastructure

Qwen3.8-2.4T

Alibaba has released Qwen3.8, a 2.4T-parameter mixture-of-experts model with 95B active parameters, long-context support, and claimed frontier-level coding and agent performance. HN focuses on its enormous serving requirements, missing vision support, quantization trade-offs, licensing, and implications for open model competition.

HN Discussion
12 Aug 2026
InfrastructureSafety and policy

Someone is running mass vulnerability scans, spoofing AI bots like ClaudeBot

A coordinated scanning campaign is impersonating AI bots such as ClaudeBot while probing for credentials and configuration files used by AI coding tools. HN commenters compare it with ordinary internet-wide scanning and discuss bot authentication, ASN filtering, and the risks of exposing agent secrets.

HN Discussion
12 Aug 2026
AgentsOpen sourceInfrastructureAI applications

My Agent Setup

A founder details a six-agent AI staff running on a small VPS, coordinated through open-source Buzz and Nostr, with agents handling development, operations, research, marketing, and administration. HN discusses the setup’s security, privacy, cost, reliability, and disappointing early ROI.

HN Discussion
12 Aug 2026
ModelsInfrastructureAI applications

Automatic1111 for Apple metal, 40% speed up sd1.5

An Automatic1111-based Stable Diffusion 1.5 implementation reportedly achieves a 40% speedup on Apple Metal, including LoRA workflows. HN discusses its continued usefulness, alternatives such as ComfyUI and SD.cpp, and the tradeoffs in their user interfaces.

HN Discussion
12 Aug 2026
InfrastructureBusiness and industry

Electricity Pricing in the Age of AI

A primer explains how AI data-center growth is reshaping electricity markets, from generation and transmission constraints to hedging and power-plant financing. It highlights why power availability and pricing may become decisive limits on AI capacity expansion.

HN Discussion
12 Aug 2026
InfrastructureSafety and policy

Chrome adopts what may be the best protection yet against account takeovers

Chrome is testing device-bound session credentials, which tie session authentication to a hardware-protected private key so stolen cookies cannot be reused for account takeover. HN discusses how DBSC differs from passkeys and the privacy implications of hardware-bound identity.

HN Discussion
← NewerPage 9Older →