Eieye
Eieye
FeedChatReportsAbout
ModelsAgentsCoding toolsOpen sourceResearchInfrastructureAI applicationsSafety and policyBusiness and industry
Filter stories by topic
ModelsAgentsCoding toolsOpen sourceResearchInfrastructureAI applicationsSafety and policyBusiness and industry

Updated 1 Sept, 17:26

23 Aug 2026
ModelsResearch

Mathematicians will probably become obsolete before anyone else [pdf] (2004)

A 2004 letter by Ted Kaczynski argues that mathematicians may be among the first professionals made obsolete by intelligent computers. HN debates whether modern LLMs support that prediction, while questioning their reliability and ability to produce meaningful mathematics.

HN Discussion
22 Aug 2026
ModelsAgentsResearch

NanoGPT Speedrun Frontier

A benchmark runs 18 frontier models through 153 autonomous nanoGPT optimization sessions, comparing their ability to conduct iterative experiments under time and token constraints. HN discusses the strong impact of agent harnesses, benchmark variance and contamination, and what the results reveal about autonomous AI research.

HN Discussion
22 Aug 2026
ModelsAgentsAI applications

English ↔ Claudish Translator

A playful English–Claudish translator targets Claude’s distinctive phrasing and failure modes. The discussion explores whether the style is intentional, how to suppress it, and how it affects Claude Code and agent workflows.

HN Discussion
22 Aug 2026

Built by Will Etheridge

wjeth.comwjeth@pm.me
Models
Research
Infrastructure

Why your local LLM feels dumber than it is

An investigation shows that local LLM quality can change with quantization, chat templates, sampling, KV-cache precision, attention algorithms, reduction order, and GPU-specific execution—not just model weights. HN commenters add practical diagnosis and tuning advice, especially for Qwen, Ollama, llama.cpp, and Apple hardware.

HN Discussion
22 Aug 2026
ModelsCoding toolsBusiness and industry

Anthropic appears to be A/B testing reduced effort levels in Claude Code

Anthropic is server-side A/B testing a different numerical mapping for Claude Code’s reasoning-effort settings, while saying the selected effort and model performance are unchanged. HN commenters debate reports of degraded Claude quality, opaque routing, token economics, and the trust implications of testing paying users.

HN Discussion
22 Aug 2026
ModelsCoding toolsBusiness and industry

GPT 5.6 Sol 20% price reduction

OpenAI is temporarily cutting GPT-5.6 Sol API prices by 20% on input and 33% on output, with higher rates for very long contexts. HN discusses whether the move reflects efficiency and competition, and compares Sol’s coding performance and value with Claude, DeepSeek, and other models.

HN Discussion
22 Aug 2026
ModelsBusiness and industry

OpenAI cuts developer pricing for frontier GPT-5.6 Sol model by more than 20%

OpenAI has cut developer pricing for its frontier GPT-5.6 Sol model by more than 20%, intensifying competition in the LLM API market.

HN Discussion
21 Aug 2026
ModelsOpen sourceResearchInfrastructure

Run 290B+ frontier MoE models locally on your gaming PC

FreeToken is an Apache-licensed inference engine that uses GPUs, CPUs, RAM, and VRAM together to run 290B+ open-weight MoE models on consumer PCs. Its paper and implementation focus on bandwidth-adaptive execution, expert caching, and efficient local serving.

HN Discussion
21 Aug 2026
ModelsSafety and policy

The Creation of Abulafia

An essay uses Umberto Eco’s fictional text-shuffling computer to examine how LLMs can make user-generated theories feel independently validated. It warns that memory, fluent elaboration, and apparent disagreement can reinforce epistemic feedback loops rather than provide genuine corroboration.

HN Discussion
21 Aug 2026
ModelsAgentsCoding tools

A week of using Codex more than Claude

A developer compares a week of using Codex TUI with Claude Code, finding Codex more concise, technical, and often faster, while Claude better infers intent and handles ambiguity. HN commenters debate model-versus-harness effects, overengineering, quotas, context management, and mixed-agent workflows.

HN Discussion
21 Aug 2026
ModelsAI applicationsSafety and policyBusiness and industry

Bringing the cybersecurity capabilities of Claude Mythos 5 to more defenders

Anthropic is expanding Claude Mythos 5 into defensive cybersecurity tools, including enterprise vulnerability scanning, partner integrations, and a $35M open-source security credit fund. HN focuses on the trade-off between cyber safeguards and defenders’ ability to investigate and reproduce vulnerabilities.

HN Discussion
21 Aug 2026
ModelsAI applications

LLMs are proof that Unix won

An essay argues that LLMs revive Unix’s text-stream interface by translating natural-language requests into shell commands. HN debates the limits of the Unix analogy, command-line conventions, and Linux’s role in hosting models locally.

HN Discussion
21 Aug 2026
ModelsOpen sourceInfrastructureAI applications

How we made a text-to-speech model respond in sub-50 ms

Nari Labs open-sources an optimized Qwen3-TTS implementation that delivers sub-50 ms p95 time-to-first-audio at up to 10 requests per second on one H100. The work combines scheduling, CUDA graphs, codec state caching, and streaming-focused tuning for real-time voice applications.

HN Discussion
21 Aug 2026
ModelsOpen sourceAI applications

Show HN: Public Muscriptor Instance (latest, most powerful Audio-to-MIDI model)

MuScriptor is an open model that transcribes every instrument in a recording into MIDI without requiring instrument labels. HN users report surprisingly accurate results while discussing its noncommercial license and the public instance’s rate limits.

HN Discussion
21 Aug 2026
ModelsAgentsCoding tools

Claudette: Make Claude stop talking like a BuzzFeed article

Claudette is a Claude Code skill that sends Claude’s verbose, hype-heavy responses through Gemini to produce plainer English. HN discusses whether model chaining is worthwhile versus prompts, hooks, or switching models, while comparing the writing styles of competing coding agents.

HN Discussion
21 Aug 2026
ModelsAgentsResearch

Nvidia AVO scores 100% on the ARC-AGI-3 interactive reasoning benchmark

NVIDIA’s AVO agent achieved 100% on ARC-AGI-3’s 183-level public set using Claude Opus 5, combining evolutionary search with autonomous long-horizon reasoning. HN discusses the result’s harness dependence, public-set limitation, and whether it says anything about AGI.

HN Discussion
21 Aug 2026
ModelsInfrastructureAI applicationsBusiness and industry

What Happens When the Cost of Intelligence Drops 100x

An analysis of how LLM capability is becoming dramatically cheaper, with some intelligence tiers falling 30–100x in roughly a year. The article and discussion examine commoditization, open models, agent economics, throughput, and whether cheaper inference will drive vastly more usage.

HN Discussion
21 Aug 2026
ModelsAI applications

I'm becoming AI-blind

The author argues that repeated exposure to low-effort AI prose has created “AI blindness”: verbose, jargon-heavy output is increasingly filtered as meaningless. HN discusses recognizable model tics, workplace and coding-documentation problems, and whether the perceived lack of intent reflects genuine model limitations or simply poor prompting and editing.

HN Discussion
21 Aug 2026
ModelsAgentsResearchAI applications

DeepSeek-v4-flash-vision-exp

DeepSeek has released an experimental vision variant of V4 Flash with OpenAI- and Anthropic-compatible image APIs. HN users discuss its usefulness for OCR, UI and agent feedback loops, while testing exposes resolution limits and uneven visual reasoning.

HN Discussion
21 Aug 2026
ModelsSafety and policyBusiness and industry

AI companies destroy physical books – let's scan rare books before it's too late

Anna’s Archive alleges that AI companies are buying, destructively scanning, and retaining private copies of books for model training, urging volunteers to preserve them publicly. HN debates the evidence and rarity of the books, while focusing on copyright law, proprietary training data, and whether digitization outweighs the loss of physical access.

HN Discussion
← NewerPage 6Older →