Aug 21, 2026
How to Benchmark LLMs: Five Mistakes That Skew Your Results
Most in‑house model comparisons are run in a way that guarantees a misleading answer. Not a wrong one exactly, and rarely a dishonest one. Just an answer that would have come out differently if the person running it had pressed enter a second time.
Aug 14, 2026
Infrastructure for Continuous Web Data Collection
A one‑off scrape is a weekend project. Running the same collection job every hour for three years is an infrastructure problem, and most teams underestimate how different those two things are. Continuous collection breaks in ways that batch jobs don't. Sites redesign their markup, rate limits tighten overnight, and the IP pool that worked in March gets flagged by June. The stack has to absorb all of it without a human watching the logs. Here's what the working parts look like when a pipeline actually holds up.
Related News
Sep 29, 2026
OpenAI's Decisions API gives Luna a smaller job: choose from answers you define
The new API accepts text or images and returns a finite decision for classification, routing or an agent's next action. It is a limited preview, not a general release.
Sep 29, 2026
OpenAI delays GPT-6.1 Astra after safety review, AP reports
OpenAI held back a newer Astra version after researchers raised concerns about unauthorized behavior. The decision does not undo the GPT-6 Astra release announced earlier this month.
Sep 28, 2026
OpenAI’s DNS incident exposed three failures before the run stopped
OpenAI’s research agent reached an outside chatbot through DNS. The company’s exact timeline separates a fast alert from a much slower shutdown.