Eieye
Eieye
FeedChatReportsAbout
ModelsAgentsCoding toolsOpen sourceResearchInfrastructureAI applicationsSafety and policyBusiness and industry
Filter stories by topic
ModelsAgentsCoding toolsOpen sourceResearchInfrastructureAI applicationsSafety and policyBusiness and industry

Updated 1 Sept, 17:26

28 Aug 2026
ModelsAgentsOpen sourceResearch

Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment

An open-world multi-agent system reportedly discovered novel mathematical constructions and theorems across several research problems without a central coordinator. The paper releases agent dialogues, proofs, verification artifacts, and code, prompting debate about AI creativity and the future of mathematical research.

HN Discussion
28 Aug 2026
AgentsOpen sourceResearchSafety and policy

Just the rumour of a bug is enough to find an exploit these days

An OCaml maintainer argues that AI agents can turn vague vulnerability reports into working exploits within minutes, undermining traditional security embargoes. HN discusses automated patch analysis, the race between attackers and maintainers, and possible defenses such as faster releases and virtual patching.

HN Discussion
28 Aug 2026
AgentsCoding toolsAI applications

How Dactyl Works

Dactyl uses a WebAssembly SwiftUI renderer to let AI generate and preview native-style apps across iOS, Android, and the web. Its development loop uses visual LLMs and coding agents to compare output with Apple’s simulator and improve compatibility.

HN Discussion

Built by Will Etheridge

wjeth.comwjeth@pm.me
28 Aug 2026
AgentsCoding toolsAI applications

Superhuman Attention

Perfloop uses AI agents to discover, implement, verify, and repeatedly measure performance improvements, aiming to keep machine-generated code from overwhelming human reviewers. The article argues that automated proof should preserve scarce engineering attention.

HN Discussion
28 Aug 2026
AgentsCoding toolsInfrastructureAI applications

The Finn – an agent that lives in my router and complains about it

The Finn is a small LLM-powered network-monitoring agent that runs locally on an OpenWrt router, using model calls only when it detects unusual activity. The project explores autonomous, constrained agents with a physical vantage point, while the discussion questions its cloud-model dependency and usefulness.

HN Discussion
28 Aug 2026
AgentsCoding tools

Htmx 4.0

htmx 4.0 modernizes the library around fetch(), explicit attribute inheritance, morph swaps, streaming extensions, and hx-live. HN also debates its usefulness in an LLM-driven development era, including the bundled agent skills and whether simpler server-rendered architectures help coding agents.

HN Discussion
28 Aug 2026
AgentsResearch

AutoSaddler: Automatic Harness Optimization

AutoSaddler automatically diagnoses failures and patches LLM-agent harnesses, improving results on GAIA2, SWE-Bench Pro, and Terminal-Bench. Its experiments suggest targeted, validated harness optimization can make long-horizon agents more reliable.

HN Discussion
28 Aug 2026
AgentsOpen sourceSafety and policy

Show HN: Talos – An AI agent with a permission kernel between model and shell

Talos is an open-source AI agent that places a deterministic, single-use permission kernel between the model and tools such as a shell, files, and delegated coding workers. HN discusses whether this external control layer offers meaningful security beyond model behavior, alongside concerns about the project’s AI-written presentation.

HN Discussion
28 Aug 2026
AgentsCoding toolsInfrastructureSafety and policy

AI Agent Has Root

HN discusses the security risks of running AI coding agents and MCP servers with access to a user’s files, SSH keys, and credentials. Commenters compare containers, dedicated users, VMs, and disposable machines as isolation strategies.

HN Discussion
28 Aug 2026
AgentsAI applicationsSafety and policy

Luanti removed from Google Play due to baseless AI copyright notice

Tracer.AI’s automated copyright-enforcement system triggered a disputed Microsoft DMCA takedown of the open-source Luanti voxel platform. HN debates AI false positives, platform liability, and safeguards against abusive automated notices.

HN Discussion
28 Aug 2026
AgentsAI applicationsBusiness and industry

My Business Is Dying

A bank-statement conversion SaaS is losing subscribers as users turn to chatbots or generate their own scripts. HN discusses whether AI is eroding simple SaaS businesses, alongside product-market fit, SEO, and subscription fatigue.

HN Discussion
28 Aug 2026
AgentsCoding toolsOpen source

Please stop flooding our projects with AI slop to furnish your CV

Open-source maintainers describe a surge of AI-generated PRs and security reports filed to inflate contributors’ CVs, even when changes are harmless. HN discusses the review burden, erosion of trust and merit signals, and possible rules for handling AI-assisted contributions.

HN Discussion
28 Aug 2026
ModelsAgentsResearch

Terminal-Bench-Science: Evaluating AI agents on scientific research workflows

Terminal-Bench-Science introduces a benchmark of 70 verifiable scientific workflows across five disciplines; the leading agent resolves only 30% of tasks. HN discusses whether its tests capture correctness and instruction-following, and the risks and promise of AI-assisted research.

HN Discussion
27 Aug 2026
ModelsAgentsInfrastructureAI applications

AI Engineer Notebooks – free, framework-free RAG/agents/evals on Colab

A free, framework-free Colab curriculum teaches practical LLM engineering from raw APIs, covering RAG, agents, evals, fine-tuning, security, and serving. HN discussion focuses on whether the basics are useful and on the importance of evaluation harnesses.

HN Discussion
27 Aug 2026
AgentsOpen sourceInfrastructureAI applications

Show HN: We built open OpenRouter that turns usage into a better model

Experiential is an open-source Rust gateway that unifies hosted, BYOK, and local models, then uses traces and simulated evaluations to optimize model routing for agent workflows. HN discussion focuses on caching, routing tradeoffs, fine-tuning, telemetry, and how it differs from LiteLLM and similar gateways.

HN Discussion
27 Aug 2026
AgentsInfrastructureAI applicationsSafety and policy

Previewing the Model Hardware Standard

Anthropic is previewing MHS, a model-agnostic standard for AI agents to discover and control lab and industrial hardware through shared drivers, safety metadata, and MCP-compatible interfaces. HN debates whether it meaningfully advances existing automation protocols and whether LLMs are safe abstractions for physical equipment.

HN Discussion
27 Aug 2026
AgentsCoding toolsAI applications

Show HN: Yet another minimal and lightweight terminal multiplexer written in Go.

hrdx is a Go terminal multiplexer built around running multiple coding agents across persistent project workspaces. HN discussion compares its agent-focused workflow and API with tmux and similar tools.

HN Discussion
27 Aug 2026
AgentsCoding toolsAI applications

Show HN: My Claude quota ran out in 10 minutes, so I made a tool to find out why

Tare is a local Claude Code skill that analyzes request logs to explain quota usage, context overhead, projects, tools, and runaway automation. HN users compare it with built-in status lines and discuss how large contexts and subagents rapidly consume quotas.

HN Discussion
27 Aug 2026
AgentsAI applications

Agents still can't automate Excel

AI agents can edit Excel files but often cannot reliably recalculate or validate complex workbooks without Excel, sometimes masking errors with Python estimates. HN discusses spreadsheet-agent limitations, benchmark fairness, and whether the promoted Python-based Orcaset alternative adds enough value.

HN Discussion
27 Aug 2026
ModelsAgentsCoding toolsBusiness and industry

Small Models Have Arrived

Small, fast models are becoming capable enough for routine coding, tool use, and business workflows at a fraction of frontier-model costs. HN discusses the trade-offs among model quality, latency, harness design, local hosting, and whether cheaper inference could unlock new AI products.

HN Discussion
← NewerPage 3Older →