Eieye
Eieye
FeedChatReportsAbout
ModelsAgentsCoding toolsOpen sourceResearchInfrastructureAI applicationsSafety and policyBusiness and industry
Filter stories by topic
ModelsAgentsCoding toolsOpen sourceResearchInfrastructureAI applicationsSafety and policyBusiness and industry

Updated 1 Sept, 17:26

29 Aug 2026
AgentsResearchAI applications

I accidentally turned LLM memory into program analysis

Lemmalog gives LLM agents a maintained Datalog state instead of relying on transcript retrieval, enabling provenance, retractions, and temporal updates. Its benchmarks show substantially smaller query context and stronger handling of knowledge updates, though extraction and inference remain weak.

HN Discussion
28 Aug 2026
ResearchAI applicationsSafety and policy

Identifying fake cosmetics using AI

A lab tests Gemini’s image analysis on authentic and counterfeit Rhode lip tints. It catches subtle packaging errors but also mistakes glare for defects and falsely labels a genuine product counterfeit, prompting debate about AI’s reliability and overconfidence.

HN Discussion
28 Aug 2026
ModelsResearch

Racter (1984)

Racter was a 1984 BASIC program that generated seemingly coherent English prose from rules, variables, and randomized selections. HN discusses it as an early ancestor of modern language models and conversational AI.

HN Discussion
28 Aug 2026

Built by Will Etheridge

wjeth.comwjeth@pm.me
Research
AI applications

The Analytical AI Handbook

The Analytical AI Handbook covers using foundation models to transform unstructured data into reliable decisions through classification, extraction, judging, and evaluation. It emphasizes measurable tasks, smaller models, batch processing, and production architectures.

HN Discussion
28 Aug 2026
ModelsResearch

Separating logic and language

An MIT-led study finds that severe language impairment does not prevent logical reasoning, suggesting distinct brain systems for language and logic. HN discussion explores whether this supports separating language interfaces from reasoning and control in LLM-based systems.

HN Discussion
28 Aug 2026
ModelsAgentsOpen sourceResearch

Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment

An open-world multi-agent system reportedly discovered novel mathematical constructions and theorems across several research problems without a central coordinator. The paper releases agent dialogues, proofs, verification artifacts, and code, prompting debate about AI creativity and the future of mathematical research.

HN Discussion
28 Aug 2026
AgentsOpen sourceResearchSafety and policy

Just the rumour of a bug is enough to find an exploit these days

An OCaml maintainer argues that AI agents can turn vague vulnerability reports into working exploits within minutes, undermining traditional security embargoes. HN discusses automated patch analysis, the race between attackers and maintainers, and possible defenses such as faster releases and virtual patching.

HN Discussion
28 Aug 2026
ModelsCoding toolsOpen sourceResearch

GLM-5.3 is now open-weight

Z.ai has released GLM-5.3 as an open-weight model, claiming major post-training gains in coding, long-horizon agents, and cybersecurity. HN users discuss real-world performance, hosting costs, quantization, local deployment, and the risks of its cyber capabilities.

HN Discussion
28 Aug 2026
AgentsResearch

AutoSaddler: Automatic Harness Optimization

AutoSaddler automatically diagnoses failures and patches LLM-agent harnesses, improving results on GAIA2, SWE-Bench Pro, and Terminal-Bench. Its experiments suggest targeted, validated harness optimization can make long-horizon agents more reliable.

HN Discussion
28 Aug 2026
Coding toolsResearch

Your AGENTS.md file doesn't do anything

An ETH Zurich study finds AGENTS.md context files generally do little to improve coding-agent task success while increasing inference costs by over 20%. HN users report practical gains from concise, project-specific instructions but debate the study’s methodology and conclusions.

HN Discussion
28 Aug 2026
ResearchInfrastructureAI applications

Benchmarking Vector Indexes

A reproducible harness benchmarks database vector indexes using recall-versus-throughput curves, build costs, filtering, concurrency, and churn. It emphasizes ground truth, query-plan validation, and pinned environments to prevent misleading ANN comparisons.

HN Discussion
28 Aug 2026
ModelsAgentsResearch

Terminal-Bench-Science: Evaluating AI agents on scientific research workflows

Terminal-Bench-Science introduces a benchmark of 70 verifiable scientific workflows across five disciplines; the leading agent resolves only 30% of tasks. HN discusses whether its tests capture correctness and instruction-following, and the risks and promise of AI-assisted research.

HN Discussion
27 Aug 2026
ModelsResearchInfrastructure

Benchmarking Pocket-Scale Inference

A benchmark compares small language models for pocket-scale, on-device inference. HN discusses the gap between benchmark scores and real-world usefulness, plus mobile RAM, power, and NPU constraints that still limit local LLM deployment.

HN Discussion
27 Aug 2026
Coding toolsResearch

We found a division by zero bug in FFmpeg with a vibecoded fuzzer

An AI-assisted fuzzer found a division-by-zero crash in FFmpeg, though commenters note the issue may have been discovered previously by OSS-Fuzz. The discussion focuses on whether LLM-built fuzzing tools provide meaningful novelty beyond conventional fuzzers.

HN Discussion
27 Aug 2026
ModelsResearch

Formalization of the Solution to the Hopf Problem

A Lean repository formalizes the solution to the Hopf problem. HN discussion focuses on the growing role of LLMs in producing and checking large-scale mathematical proofs, and what human understanding means when formalization is machine-assisted.

HN Discussion
27 Aug 2026
AgentsResearchAI applicationsBusiness and industry

Launch HN: Salem Robotics (YC S26) – Software for industrial inspection robots

Salem Robotics develops an AI-plus-classical-robotics software layer for hazardous industrial inspections, using existing robot hardware for tasks such as nuclear surveys and leak detection. The discussion explores the practical boundary between learned perception and explicit geometric planning in safety-critical robots.

HN Discussion
27 Aug 2026
ResearchAI applications

Engineered yeast for converting plastic and biomass compounds to food additives

Researchers engineered yeast to turn preprocessed PET-derived compounds and biomass into proteins, fats, and other food ingredients. HN discusses the process’s reliance on chemical pretreatment, uncertain economics, environmental safety, and whether consumers would accept the resulting food.

HN Discussion
27 Aug 2026
AgentsResearchInfrastructure

Needle: The benchmark your search engine can't memorize

NEEDLE is a continuously refreshed benchmark for evaluating search engines used by AI agents, aiming to prevent memorization and data leakage. HN discusses its methodology, reward-hacking risks, and the credibility challenge of a search provider benchmarking itself.

HN Discussion
27 Aug 2026
ResearchAI applications

Navigation app for people with blindness and low vision

Harvard researchers developed Mobilio, an on-device computer-vision navigation app that uses smartphone sensors and spatial audio to guide blind and low-vision users. Early tests found faster routes and fewer obstacle collisions than Google Maps plus a cane.

HN Discussion
27 Aug 2026
AgentsCoding toolsResearchAI applications

Software engineering is about managing complexity

An essay argues that AI makes code generation cheap but leaves architecture, tradeoffs, verification, and long-term ownership to engineers. HN debates whether agents will soon manage complexity themselves or remain limited by context and unreliable judgment.

HN Discussion
← NewerPage 2Older →