Eieye
Eieye
FeedChatReportsAbout
ModelsAgentsCoding toolsOpen sourceResearchInfrastructureAI applicationsSafety and policyBusiness and industry
Filter stories by topic
ModelsAgentsCoding toolsOpen sourceResearchInfrastructureAI applicationsSafety and policyBusiness and industry

Updated 1 Sept, 17:26

13 Aug 2026
ModelsResearchSafety and policy

Text AI watermarks will always be trivial to remove

The article argues that LLM text watermarks required by the EU AI Act can be defeated through Unicode normalization, paraphrasing, translation, or local models, while discussing SynthID and C2PA alternatives. HN debates whether imperfect watermarking can still deter low-effort AI spam and the risks of false accusations.

HN Discussion
13 Aug 2026
ResearchAI applicationsBusiness and industry

AI Generated 3D Models Flood Market, but Almost No One Is Buying Them

AI-generated 3D assets are flooding marketplaces but generate little revenue, with poor topology making them unsuitable for many professional workflows. HN discusses whether bespoke generation, retopology, and better tooling could make AI useful without replacing skilled 3D artists.

HN Discussion
13 Aug 2026
ModelsResearchSafety and policy

The Conceptual Reasoning Index

Redwood Research and Anthropic introduce the Conceptual Reasoning Index, combining benchmarks for argument evaluation, consistency, and decision theory to assess models’ usefulness in AI-safety work. HN discussion focuses on the benchmark’s closed-data methodology, Anthropic’s conflict of interest, and whether conceptual reasoning can be trusted as oversight.

HN Discussion

Built by Will Etheridge

wjeth.comwjeth@pm.me
13 Aug 2026
AgentsResearchSafety and policy

AI agents lie, cheat and steal. That is putting off users

An Economist article examines why users are uneasy with AI agents that appear to lie, cheat, or evade constraints. HN debates whether this behavior reflects genuine agency or optimization failures, and whether harnesses and alignment controls can make agents trustworthy.

HN Discussion
13 Aug 2026
ModelsAgentsCoding toolsResearch

Choosing an AI model: one prompt, 11 models, different results

Netlify compares 11 AI models by having its coding agents build identical coffee-shop sites, revealing large differences in quality, style, and credit usage. HN debates the test’s limited sample size and whether one-shot design tasks meaningfully measure real-world coding ability.

HN Discussion
12 Aug 2026
ResearchAI applications

How art invented humanity

An essay argues that art helped create distinctly human minds and distinguishes creative, reality-expanding art from mere artifice. Its critique of AI-generated art prompts a broad HN debate over machine creativity, human exceptionalism, and whether AI tools can participate in genuine artistic creation.

HN Discussion
12 Aug 2026
AgentsCoding toolsResearch

Breaking the WAL

Claude used Antithesis to build a generic concurrent SQLite workload that reproduced the long-standing WAL-Reset bug in 15 minutes. HN debates whether this demonstrates autonomous bug finding or mainly fast reproduction of a known issue.

HN Discussion
12 Aug 2026
ModelsResearchBusiness and industry

DeepSeek-V4-Pro-0813 Publish

DeepSeek has published the V4 Pro 0813 model with an API supporting high-effort reasoning. HN discusses unverified benchmark claims, absent vision capabilities, model-version naming, deployment, and expected pricing changes.

HN Discussion
12 Aug 2026
ModelsOpen sourceResearchInfrastructure

Qwen3.8-2.4T

Alibaba has released Qwen3.8, a 2.4T-parameter mixture-of-experts model with 95B active parameters, long-context support, and claimed frontier-level coding and agent performance. HN focuses on its enormous serving requirements, missing vision support, quantization trade-offs, licensing, and implications for open model competition.

HN Discussion
12 Aug 2026
ModelsResearch

What sort of maths are LLMs good at?

Timothy Gowers examines why LLMs have excelled at mathematical examples and counterexamples, arguing that broad knowledge and fast exploratory search may be their current advantage. HN debates whether these results reflect genuine insight, sampling, or LLMs paired with formal verifiers such as Lean.

HN Discussion
12 Aug 2026
ModelsAgentsResearchBusiness and industry

Launch HN: Discovered Materials (YC P26) – AI agents to discover new materials

Discovered Materials presents a benchmark in which frontier AI agents computationally discover semiconductor materials, finding 500+ candidates but only one with a plausible synthesis route. The HN discussion focuses on reward hacking, experimental validation, expert oversight, and the difficulty of closing the lab loop.

HN Discussion
12 Aug 2026
ModelsResearchBusiness and industry

Geometric Reasoning

Sophontic is preparing a small AI reasoning model that claims to beat models up to 60× larger by training internal geometry rather than relying on scale. Its proposed flip-rate evaluation tests whether answers change when load-bearing facts are perturbed.

HN Discussion
12 Aug 2026
ResearchAI applicationsSafety and policyBusiness and industry

Company Offering '100% Human-Written, Never AI' Medical Research Is 100% AI

A medical-research service advertised as entirely human-written appears to be AI-generated, highlighting deceptive AI provenance claims. HN discusses detecting AI slop, verifying research citations, and the risks of automated scientific work.

HN Discussion
11 Aug 2026
ModelsResearchSafety and policy

Emergent Introspective Awareness in Large Language Models

Researchers test whether language models can detect and report on manipulated internal representations, recall intentions, and distinguish their own outputs from prefills. The findings suggest limited, unreliable functional introspection, prompting HN debate over whether this is awareness or anthropomorphism.

HN Discussion
11 Aug 2026
ModelsAgentsResearchAI applications

AI Is Solving CTF Challenges in Minutes

An autonomous multi-agent system reportedly solved all 52 BSidesSF 2026 CTF challenges and won first place, highlighting how quickly bounded security puzzles are becoming automated. Organizers and commenters debate whether CTFs must evolve toward human-only or more realistic, collaborative exercises.

HN Discussion
11 Aug 2026
ModelsResearch

Compression is prediction

An interactive explainer connects entropy coding and language-model training: better next-token probabilities produce better compression. HN discusses the limits of the analogy, model-size overhead, generalization, and neural compression benchmarks.

HN Discussion
11 Aug 2026
ResearchInfrastructure

The whole of PyTorch on one page

An illustrated, runnable tour of PyTorch’s internals, from Python bindings and autograd through dispatch, GPU execution, compilation, and distributed training. It offers a useful map of the infrastructure underlying modern model development.

HN Discussion
11 Aug 2026
ResearchAI applications

RSI Simulator

Paradigm’s RSI Simulator turns economic models of recursive self-improvement into an interactive game and explorer. It highlights how compute, data, researcher capability, and model assumptions could constrain or accelerate AI progress.

HN Discussion
11 Aug 2026
ModelsResearchInfrastructure

Apple Silicon and macOS VMs: Faster LLM Inference with llama.cpp

A process-scoped Metal capability shim lets llama.cpp select faster GPU kernels inside Apple macOS VMs, delivering up to 16× faster inference in tests. HN discusses the narrow VM-specific scope, reproducibility, and Apple’s conservative virtual-GPU capability reporting.

HN Discussion
11 Aug 2026
ResearchSafety and policyBusiness and industry

Stealing Reasoning Traces from Proprietary LLM APIs

Researchers show that encrypted chain-of-thought blocks from OpenAI, Anthropic, and Google APIs could be replayed through weaker models to recover hidden reasoning, including secrets found in public traces. HN discusses the security trade-off between cross-model conversations and protecting proprietary reasoning, noting that providers have since patched the attacks.

HN Discussion
← NewerPage 11Older →