Eieye
Eieye
FeedChatReportsAbout
ModelsAgentsCoding toolsOpen sourceResearchInfrastructureAI applicationsSafety and policyBusiness and industry
Filter stories by topic
ModelsAgentsCoding toolsOpen sourceResearchInfrastructureAI applicationsSafety and policyBusiness and industry

Updated 1 Sept, 17:26

24 Aug 2026
ModelsCoding toolsInfrastructure

Claude Is Down?

Claude experienced a widespread, temporary outage affecting Opus, Fable, and many Claude Code sessions, with users reporting 529 overload errors across multiple regions. The discussion highlights growing reliance on AI coding services and the need for provider or local-model fallbacks.

HN Discussion
23 Aug 2026
ModelsResearchInfrastructure

Etched Sohu vs. Nvidia: Transformer ASIC vs. GPU (2026)

Etched’s Sohu is a transformer-only inference ASIC claiming dramatically higher Llama throughput than GPUs, but it sacrifices programmability and remains independently unverified. HN discusses whether attention hardware is really the bottleneck and how well the chip can support evolving models and workloads.

HN Discussion
23 Aug 2026
AgentsCoding toolsInfrastructure

AI and Infrastructure Engineering

An infrastructure engineer argues that AI is moving automation up another abstraction layer, generating Terraform and Helm while humans retain architectural judgment. HN discusses the productivity gains, risks of losing foundational skills, and whether LLMs can reliably reason about complex infrastructure.

HN Discussion

Built by Will Etheridge

wjeth.comwjeth@pm.me
23 Aug 2026
AgentsOpen sourceInfrastructure

What Is a Harness?

An accessible explanation of AI agent harnesses—the prompts, tools, control loops, and translation layers that turn models into usable agents. HN discusses open-source alternatives such as Pi and Goose, customization, guardrails, handoff, and whether harnesses are strategic or commodity software.

HN Discussion
23 Aug 2026
Coding toolsResearchInfrastructureAI applications

JIT Compiling Code in 5μs

An AI-assisted copy-and-patch JIT compiler generates ARM64 machine code in about 5μs, letting pgrust compile every SQL query and approach handwritten performance. HN debates whether AI meaningfully lowers the barrier to complex compiler work, alongside JIT security and optimization trade-offs.

HN Discussion
23 Aug 2026
ResearchInfrastructure

AI Chip Architectures

A deep survey compares GPUs, TPUs, Trainium, Cerebras, and Groq across compute, memory, interconnects, software, and scaling for modern AI workloads. HN discussion extends the analysis to power consumption and whether analog or neuromorphic designs could outperform today’s LLM accelerators.

HN Discussion
22 Aug 2026
ModelsResearchInfrastructure

Why your local LLM feels dumber than it is

An investigation shows that local LLM quality can change with quantization, chat templates, sampling, KV-cache precision, attention algorithms, reduction order, and GPU-specific execution—not just model weights. HN commenters add practical diagnosis and tuning advice, especially for Qwen, Ollama, llama.cpp, and Apple hardware.

HN Discussion
22 Aug 2026
InfrastructureSafety and policyBusiness and industry

Anthropic IPO filing will show AI backlash as a risk factor, sources say

Anthropic is preparing an IPO that could value the Claude maker near $2 trillion, while warning that public opposition to AI and data-center construction could threaten growth. HN debates profitability, compute costs, political backlash, and the industry’s broader social risks.

HN Discussion
22 Aug 2026
InfrastructureBusiness and industry

The war on data centres is a bit fake

HN debates whether opposition to data centres is manufactured or a legitimate response to AI-driven expansion. The discussion focuses on power demand, noisy temporary generators, grid costs, and who should pay for infrastructure.

HN Discussion
22 Aug 2026
AgentsOpen sourceInfrastructureSafety and policy

New MCP Roadmap

The MCP roadmap prioritizes agentic messaging, unified HTTP transport, progressive tool discovery, SDK improvements, and enterprise agent identity and authorization. HN debates whether MCP’s growing complexity is justified versus simpler OpenAPI, REST, CLI, or code-mode approaches.

HN Discussion
22 Aug 2026
InfrastructureAI applications

Embedded AI

A sample chapter from Embedded AI teaches deployment of wake-word detectors, noise suppression, person detection, and other models on Arduino- and Raspberry Pi-class hardware. HN discusses the book’s value versus ChatGPT, possible AI-generated content, and whether it covers practical quantization.

HN Discussion
22 Aug 2026
AgentsInfrastructureAI applications

Show HN: OzBrain, a shared brain for knowledge between agents and your team

OzBrain is a hosted shared knowledge layer that lets multiple AI agents and teammates read, write, version, and audit a common brain. HN discusses whether it offers enough over Git, Obsidian, and LLM-Wiki, while highlighting retrieval quality, summarization drift, and conflict resolution.

HN Discussion
21 Aug 2026
ModelsOpen sourceResearchInfrastructure

Run 290B+ frontier MoE models locally on your gaming PC

FreeToken is an Apache-licensed inference engine that uses GPUs, CPUs, RAM, and VRAM together to run 290B+ open-weight MoE models on consumer PCs. Its paper and implementation focus on bandwidth-adaptive execution, expert caching, and efficient local serving.

HN Discussion
21 Aug 2026
Open sourceInfrastructure

OTel isn’t going well

An analysis argues that OpenTelemetry’s broad scope, strict stability commitments, and thin maintainer bench make the project slow and difficult to use. HN users debate its complexity and overhead while acknowledging its importance as a vendor-neutral observability standard.

HN Discussion
21 Aug 2026
AgentsCoding toolsInfrastructureSafety and policy

Building an (almost) fully self-hosted, sandboxed, agentic software factory

An isolated homelab runs Hermes with Codex, Forgejo, Coolify, and Firecrawl to turn one prompt into tested, deployed software. HN discusses local-model trade-offs, verification bottlenecks, and how much autonomy agents should receive.

HN Discussion
21 Aug 2026
ModelsOpen sourceInfrastructureAI applications

How we made a text-to-speech model respond in sub-50 ms

Nari Labs open-sources an optimized Qwen3-TTS implementation that delivers sub-50 ms p95 time-to-first-audio at up to 10 requests per second on one H100. The work combines scheduling, CUDA graphs, codec state caching, and streaming-focused tuning for real-time voice applications.

HN Discussion
21 Aug 2026
ModelsInfrastructureAI applicationsBusiness and industry

What Happens When the Cost of Intelligence Drops 100x

An analysis of how LLM capability is becoming dramatically cheaper, with some intelligence tiers falling 30–100x in roughly a year. The article and discussion examine commoditization, open models, agent economics, throughput, and whether cheaper inference will drive vastly more usage.

HN Discussion
21 Aug 2026
Coding toolsInfrastructureBusiness and industry

Codex on AWS bedrock bug causing 10x charges

A Codex CLI bug on Amazon Bedrock appears to disable effective prompt-cache reads while generating expensive cache writes, producing reported costs around 10× higher. Users discuss usage spikes, release-note transparency, and disabling web search as a workaround.

HN Discussion
21 Aug 2026
ModelsInfrastructureSafety and policy

Ox Alpha

OpenRouter’s anonymous Ox Alpha model became a large-scale community investigation, with testing suggesting it is ZAI’s GLM-5.3 Flash. Discussion covers its coding and reasoning quality, censorship fingerprints, uncertain provider identity, and privacy trade-offs.

HN Discussion
20 Aug 2026
InfrastructureBusiness and industry

Protesters haul a guillotine to city council meeting about an AI data center

Protesters brought a guillotine to a city council meeting over an AI data center. HN discussion focuses on whether data centers raise utility costs and strain local infrastructure.

HN Discussion
← NewerPage 5Older →