DeepSeek peak/off-peak pricing update
DeepSeek launches its V4-Pro model with stronger agent workflows, flexible reasoning effort, and OpenAI Responses API support. API pricing will vary by demand, with off-peak rates set 50% below peak.
DeepSeek launches its V4-Pro model with stronger agent workflows, flexible reasoning effort, and OpenAI Responses API support. API pricing will vary by demand, with off-peak rates set 50% below peak.
Z.ai’s GLM-5.3 claims major coding and cybersecurity gains over GLM-5.2, driven entirely by post-training, with weights promised in two weeks. HN users debate its performance against closed frontier models, practical coding use, hardware demands, vulnerability discovery, and the risks of releasing unrestricted cyber capabilities.
Terence Tao presents a human-readable digestion of an AI-generated, Lean-verified proof resolving Sendov’s conjecture, while explaining the underlying mathematics and its limits. HN debates whether AI changes mathematical understanding, discovery, and the value of human-led proof development.
Qwen releases Qwen3.8-27B, a 27B open vision-language model supporting configurable reasoning, 262K-native context, video understanding, and agentic coding. HN users discuss its impressive benchmark gains and the challenge of running it on consumer hardware.
Comma.ai’s Chestnut is an open-source-firmware USB4 eGPU dock designed to run substantially larger openpilot driving models in cars. HN discusses its tinygrad integration, hardware limitations, cost, and safety fallback behavior.
HN users compare monthly spending on Claude, ChatGPT, Gemini, coding assistants, APIs, and local models. The discussion highlights wide cost variation, from free and local setups to hundreds or thousands of dollars, along with debates over subscription value and usage limits.
OpenAI and Cerebras are previewing Ultrafast, a GPT-5.6 Sol API tier delivering up to 750 output tokens per second. HN debates its benchmark claims, likely premium pricing, Cerebras’s wafer-scale architecture, and how low-latency inference could change agentic coding and real-time applications.
Twitch will let Amazon use creators’ livestreams to train generative AI models unless they opt out, prompting backlash over consent and transparency. HN commenters focus on the deliberately non-consensual default and Twitch’s uncertainty about whether training has already occurred.
Google’s Gemini 3.7 Flash targets coding and agent workflows with reported benchmark gains, faster responses, and introductory API pricing of $0.75 per million input tokens. HN users debate its real-world quality, strong multimodal performance, competitiveness against cheaper models, and Google’s confusing API and billing experience.
Mistral OCR 4.1 is discussed as a faster, specialized document-understanding model with layout and bounding-box extraction. HN users debate its accuracy, hallucinations, privacy, and steep pricing versus local and competing OCR tools.
Google introduces Gemini 3.7 Flash, a cheaper model aimed at coding, complex workflows, and autonomous agents. It claims major benchmark gains over 3.6 Flash and adds updated cyber and CBRN safeguards.
OpenAI is previewing GPT-5.6 Sol’s Ultrafast mode, promising speeds up to 14× faster through Cerebras hardware. HN discussion focuses on the infrastructure partnership and likely price premium.
A developer builds a four-GPU home inference server from used AMD hardware, custom cooling, and salvaged parts. HN discusses ROCm support, local model performance, and whether self-hosting is worth the cost versus cloud APIs.
Samsung is using Claude to generate chip-verification code and test environments, reportedly cutting some tasks from weeks to days. HN discussion questions the headline while examining productivity, source accuracy, and how generated code can be validated.
DeepSeek is raising V4 API prices by up to 1000% for cached input and adding peak/off-peak rates. HN discusses the misleading headline and notes DeepSeek remains cheaper than competing frontier models, depending on caching and usage.
The article argues that LLM text watermarks required by the EU AI Act can be defeated through Unicode normalization, paraphrasing, translation, or local models, while discussing SynthID and C2PA alternatives. HN debates whether imperfect watermarking can still deter low-effort AI spam and the risks of false accusations.
Redwood Research and Anthropic introduce the Conceptual Reasoning Index, combining benchmarks for argument evaluation, consistency, and decision theory to assess models’ usefulness in AI-safety work. HN discussion focuses on the benchmark’s closed-data methodology, Anthropic’s conflict of interest, and whether conceptual reasoning can be trusted as oversight.
An open-source personal search engine uses a small local language model to summarize and categorize 560,000 domains for about $10 in GPU time. The build report and discussion examine crawl steering, model-generated taxonomy problems, local inference costs, and AI-assisted content concerns.
Netlify compares 11 AI models by having its coding agents build identical coffee-shop sites, revealing large differences in quality, style, and credit usage. HN debates the test’s limited sample size and whether one-shot design tasks meaningfully measure real-world coding ability.
DeepSeek launched V4 Pro and Flash with agent-focused upgrades, Responses API support, and peak/off-peak pricing. HN users compare the steep cache and output increases with competing models, third-party providers, and self-hosting options.