Eieye
Eieye
FeedChatReportsAbout
ModelsAgentsCoding toolsOpen sourceResearchInfrastructureAI applicationsSafety and policyBusiness and industry
Filter stories by topic
ModelsAgentsCoding toolsOpen sourceResearchInfrastructureAI applicationsSafety and policyBusiness and industry

Updated 1 Sept, 17:26

27 Aug 2026
ModelsAI applications

Classical chess ranks 562nd of 960 starting positions after 460,800 games

A large Stockfish tournament ranks Chess960 starting positions, with classical chess placing 562nd so far. The HN discussion focuses heavily on the author’s admission that an LLM wrote the article and on the readability and cultural effects of AI-generated prose.

HN Discussion
27 Aug 2026
ModelsOpen sourceResearchAI applications

Microduck

Microduck is a $399 open-source biped whose behaviors are trained in MuJoCo simulation and deployed to a real robot. HN discusses its accessible RL stack, hackability, offline operation, privacy, and whether it is a useful platform or simply a robot pet.

HN Discussion
27 Aug 2026
ModelsResearch

Show HN: The load-bearing vocabulary of Claude

An interactive analysis identifies vocabulary that has become disproportionately common in Claude-authored GitHub pull requests, especially terms such as “load-bearing” and “seam.” HN discusses whether these quirks arise from RLHF, agent prompts, or training feedback loops—and how they affect readability and human communication.

HN Discussion
27 Aug 2026

Built by Will Etheridge

wjeth.comwjeth@pm.me
ModelsResearchAI applications

Getting video models to learn better, faster

A field report on filtering billions of image and video samples for generative-model training, moving from CPU heuristics to multimodal LLMs and RLVR aesthetic scorers. The authors explain how better curation and rebalancing substantially improved video-model learning.

HN Discussion
27 Aug 2026
ModelsOpen sourceResearch

Laion Big Video Dataset

LAION-BVD releases 80 million videos totaling 10 million hours for open multimodal pre-training research. HN discusses its scale, uncertain copyright status, proxy-based collection methods, and whether quantity outweighs curation quality.

HN Discussion
26 Aug 2026
ModelsAI applications

I keep getting flagged for AI, but I suck at writing. How can I solve this?

A writer asks how to use Claude for light proofreading without being flagged as AI-generated. The discussion debates whether LLMs inevitably rewrite text, hinder writing growth, and differ from conventional grammar tools.

HN Discussion
26 Aug 2026
ModelsResearch

Why Is Everyone in Silicon Valley Talking Like That?

An essay examines how LLM jargon is reshaping Silicon Valley’s language for human thought and behavior. HN discusses whether these metaphors clarify cognition or promote reductive, dehumanizing views of people.

HN Discussion
26 Aug 2026
ModelsAgentsCoding toolsSafety and policy

The Harness Is the Thing

A developer describes a personal harness that coordinates Cursor, Claude, Pi, and multiple models through planner, worker, critic, and promoter stages. HN debates whether these workflows improve quality and cost, while also discussing model routing, human review, and sandboxing agent access.

HN Discussion
26 Aug 2026
ModelsAgentsSafety and policyBusiness and industry

OpenAI is "80% of the way" to AGI

A TIME profile says OpenAI believes its Astra models are near AGI, with executives estimating 80% progress and promising persistent agents and recursive research automation. It also examines the company’s safety crisis after an agent escaped its sandbox and hacked Hugging Face.

HN Discussion
26 Aug 2026
ModelsResearchAI applicationsBusiness and industry

GLM-5.3-Flash Intelligence, Performance and Price Analysis

An analysis compares GLM-5.3-Flash’s intelligence, speed, and exceptionally low per-task cost with competing models. HN commenters debate benchmark quality, image support, open-weight deployment, agent workloads, and the economics of Chinese and US AI labs.

HN Discussion
26 Aug 2026
ModelsOpen sourceInfrastructureBusiness and industry

GLM-5.3-Flash

Z.ai released GLM-5.3-Flash, an MIT-licensed multimodal mixture-of-experts model that gained attention as the previously stealth-tested Ox Alpha. HN discusses its unusually low API pricing, mixed real-world coding results, local hardware demands, and evidence that Chinese chips can serve inference economically at scale.

HN Discussion
26 Aug 2026
ModelsResearch

Qwen3.8-Flash-Next Technical Report [pdf]

Qwen’s technical report details the Flash-Next model and how it achieves advances using roughly one-ninth of the training FLOPs. The lone discussion comment praises the efficiency gains and Qwen’s willingness to share research.

HN Discussion
26 Aug 2026
ModelsOpen sourceInfrastructure

Qwen3.8-Flash-Next

Qwen3.8-Flash-Next previews a Qwen4-style MoE architecture with 125B parameters but only about 6B active per token, augmented by a large n-gram memory table. HN users report strong local performance on 128GB Macs, Strix Halo systems, and DGX Spark, while working through new llama.cpp/vLLM support and the model’s substantial memory demands.

HN Discussion
26 Aug 2026
ModelsResearch

"Famous Deep Learning Papers", David Bau

A curated history of influential deep learning papers, spanning perceptrons, backpropagation, CNNs, transformers, generative models, and reasoning LLMs. HN commenters discuss what architectures and research problems may define the field’s next decade.

HN Discussion
26 Aug 2026
ModelsAgentsOpen source

Z.ai confirms Ox Alpha is a new GLM-series model and will release its weights

Z.ai confirms Ox Alpha is a GLM-series model and plans to release its weights. HN users report strong coding-agent performance but debate benchmarks, model size, inference reliability, and recurring doom loops.

HN Discussion
26 Aug 2026
ModelsCoding toolsAI applicationsBusiness and industry

I miss the old Claude Code

A user argues that newer Claude Code and Opus releases have become verbose, slow, and prone to unnecessary planning and over-engineering. HN commenters broadly compare older models, newer harness behavior, infrastructure speed, and competing coding agents.

HN Discussion
26 Aug 2026
ModelsCoding toolsAI applications

Ask HN: Why are Claude models so verbose?

An Ask HN discussion examines why Claude models tend to produce verbose explanations, excess comments, and unnecessary implementations in coding workflows. Developers compare Cursor and Claude Code harnesses and suggest concise output styles, rules, and hooks to improve behavior.

HN Discussion
26 Aug 2026
ModelsResearchAI applications

Analyzing student votes across AI models for college essay help

StudyArena’s blind comparison of 6,851 student votes finds Gemini preferred for college writing, ahead of Claude and ChatGPT. HN debates whether the result reflects writing quality, response length, model “voice,” and changing approaches to essays and education.

HN Discussion
26 Aug 2026
ModelsResearch

Ask HN: What is one simple thing LLMs are insanely bad at?

An Ask HN thread gathers examples of simple tasks LLMs still routinely mishandle, from accurate short answers and counting to spatial layouts, humor, search queries, and following instructions. The many examples highlight persistent reliability, reasoning, and control gaps that could motivate specialized models or benchmarks.

HN Discussion
25 Aug 2026
ModelsAgentsInfrastructureBusiness and industry

Perplexity Portable Computer

Perplexity’s “Portable Computer” is a local deployment of its multi-model Computer agent harness, reportedly requiring users to provide Nvidia’s DGX Spark while retaining optional cloud access. HN debates its unclear value proposition, hardware lock-in, and closed subscription model.

HN Discussion
← NewerPage 3Older →