The current strict response schema
This is the same JSON schema object sent with current generation requests, not a prose approximation or a historical response. It contains no article text or private configuration. Open the standalone JSON schema.
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"type": "object",
"properties": {
"digest": {
"type": "object",
"properties": {
"description": {
"type": "string"
},
"lede": {
"type": [
"string",
"null"
]
},
"sections": {
"type": "array",
"items": {
"type": "object",
"properties": {
"heading": {
"type": "string"
},
"items": {
"type": "array",
"items": {
"type": "object",
"properties": {
"item_id": {
"type": "string"
},
"paragraph": {
"type": "string"
},
"tags": {
"type": "array",
"items": {
"type": "string"
}
}
},
"required": [
"item_id",
"paragraph",
"tags"
],
"additionalProperties": false
}
},
"lede": {
"type": [
"string",
"null"
]
},
"topic_id": {
"type": "string",
"enum": [
"ai",
"software",
"pharma",
"healthtech",
"economy"
]
}
},
"required": [
"heading",
"items",
"lede",
"topic_id"
],
"additionalProperties": false
}
},
"tags": {
"type": "array",
"items": {
"type": "string"
}
},
"title": {
"type": "string"
}
},
"required": [
"description",
"lede",
"sections",
"tags",
"title"
],
"additionalProperties": false
}
},
"required": [
"digest"
],
"additionalProperties": false
} 1. Acquisition
The run starts with the hand-maintained source roster. It prefers an explicitly configured RSS or Atom feed. Where a publisher has no feed, a source may use a constrained sitemap, feed discovery, a YouTube channel feed, or extraction from the public page itself. Each source contributes at most twelve recent items, which keeps prolific publishers from overwhelming the week.
Requests use an identifying bot user agent, compressed responses, a sixteen-request global concurrency ceiling, and a separate 1.5-second start interval for each host. The fetcher honors stricter crawl-delay rules for HTML access and stops contacting a host for the run after a 429. Explicitly published feed URLs are treated as the publisher's feed endpoint; robots rules are checked before HTML discovery and extraction. Requests reject private and special-use addresses both in URLs and in DNS answers. The connection uses a validated address without resolving the name again, while HTTPS still verifies the original hostname. Every redirect repeats these checks, and decoded responses are capped at 15 MiB. This is an application-level guard, not an operating-system network sandbox. Fetched text remains untrusted input.
A collection-only job captures feeds daily at 23:00 UTC without model credentials or publication. Monday's run uses the retained rolling seven-day collection, reapplies the current roster and twelve-item source ceiling, and removes duplicates. This helps when entries disappear from short feeds, but daily polling, cache eviction, and the source ceiling can still lose reporting. The retained collection stays private and is not a permanent archive.
Each current-day collector invocation can attempt up to 20 additional public article reads from its 120 highest-ranked non-video cluster representatives, two at a time. Missing evidence has priority over expanding short usable summaries. These reads obey robots rules at every redirect, defer when robots is unavailable, and reject recognized paywalls, non-article responses, and mismatched titles. Existing excerpts are reused when the item identity, URL, and title still match. A manual retry can repeat the bounded acquisition work. Historical, mocked, and recovery runs do not make these extra reads. Recovered text can restore eligibility, but does not change the original summary, numerical score, or source tier.
The date window is enforced after every acquisition path. URLs are canonicalized, blocked domains are removed, duplicate items are collapsed, and summaries are bounded before anything reaches editorial logic. A source can fail or return no current entries without being silently removed from the roster.
2. Clustering coverage
Before ranking, the pipeline groups likely duplicates across the full fetched set. Two items can join a cluster when their canonical URLs match or when their titles share at least two distinctive terms and reach a Jaccard similarity of 0.5. Every new item must match every existing member, which prevents a chain of vaguely related headlines from becoming one invented story.
This is intentionally high precision and low recall. It catches reposts and close syndication, but differently worded reporting about the same event may remain separate. The cluster representative is the item from the stronger source tier, then the newer item when tiers tie.
3. Deterministic triage
Every candidate receives the same mechanical score before a language model sees the slate. Each additional source name in its cluster adds 30 points. The source tier adds 24 points for tier 0, 16 for tier 1, 8 for tier 2, and none for tier 3. Hacker News points add up to 15 more when that signal exists. Ties resolve by source tier, recency, and a stable item identifier.
Only one representative from a cluster can enter the feature bucket. An item needs a summary beyond a bare title, category label, or teaser, or retained source text, to qualify. Video metadata stays in discovery and cannot enter the writing pass. The first 120 qualifying representatives become feature candidates, using the same limit as the writer rather than a separate hidden cutoff. The next 140 items form an appendix pool; the rest are retained only as telemetry. If the feature pool is unusually small, the strongest eligible appendix items can be promoted to meet the minimum viable run. Metadata-only inputs are never promoted just to meet that minimum.
Tiers are a source-level trust prior, not a verdict on an article and not a publication quota. Tier 0 contains the highest-priority voices to avoid missing; tier 1 is the premium pool; tier 2 broadens discovery; tier 3 is fallback and perspective. A lower-tier item can outrank a higher-tier one through wider coverage or attention, and the writer is instructed to prefer primary evidence and original reporting.
4. Composition
The writer receives at most 120 top-scoring items, including their titles, bounded summaries, dates, topics, source names, URLs, and tiers. When collection retained more source text, the writer receives an excerpt of up to 4,000 characters; the original 800-character summary remains the ranking input. Selection is score-only, without reserved topic slots. Composition runs without network access or provider credentials. A trusted parent brokers at most two model requests within one six-minute generation deadline; the worker has a separate seven-minute outer deadline for completion and evidence export. The writer chooses consequential developments, avoids repeating one event, organizes them into the fixed topic taxonomy, and writes explanations from supplied evidence.
The prompt asks for concrete specifics, attribution of company and study claims, topic breadth, and restraint around engagement metrics. It discourages routine product updates and asks the writer not to let one prolific source define a section. There is no fixed number of published stories, words per item, or allocation of longer explanations. Importance, explanatory value, and available evidence determine depth; busy weeks can be longer. Smaller supported developments can appear as one- or two-sentence briefs after fuller items, without a brief quota or filler. The five domains remain AI, software engineering, pharma and biotech technology, healthtech, and economy. Software is not limited to AI coding. Life-science coverage emphasizes technology for discovery, development, trials, and manufacturing, rather than indiscriminate basic-science updates. Video metadata alone cannot establish what was said or shown. A structured response schema limits topic identifiers and requires citation item IDs. The production default is OpenAI, while manual comparison runs can use the same pipeline with Anthropic or Google; there is no silent provider fallback.
5. Citation and prose checks
Every published item must cite an item the writer was actually given. Unknown IDs, duplicate citations, incomplete structured responses, and empty paragraphs fail validation or trigger a bounded retry. Source titles, URLs, and publisher names come directly from the selected inputs. A final assertion rejects unknown or repeated citations, duplicate URLs, or changed source metadata instead of guessing a match or publishing a reduced issue. A sanitizer then applies the public voice rules and checks banned phrases and punctuation.
These controls establish provenance, not truth. A valid citation proves that the linked fetched item was in the model's input; it does not prove that every sentence follows from the source, that the source is correct, or that two outlets reported independently. Feed summaries can also be abbreviated or publisher-written. The policy therefore requires each paragraph to use only its own cited input for factual claims and to attribute claims, but the final issue remains an automated synthesis rather than fact-checked reporting.
6. Automation and publishing
The pipeline is TypeScript running under permission-bounded Deno. GitHub Actions starts it each Monday at 11:00 UTC, with manual backfill and dry-run modes available. It writes Markdoc content for the Astro site and a machine-readable process record. An independent publisher accepts bounded inert JSON, validates its content, and renders only the selected issue and evidence record at fixed paths. Scheduled runs cannot overwrite an existing issue. It rejects images, raw HTML, dynamic expressions, unapproved components and links outside the evidence record before building the site. Historical evidence compaction is computed by the publisher itself; it cannot arrive as article edits in the generated patch. Successful builds are committed to the main branch.
Cloudflare Workers Builds deploys that commit. The workflow polls the Cloudflare Workers build check for that commit and opens a labeled GitHub issue when generation or verification fails; a later successful run closes the failure issue. Credentials remain workflow secrets and are never written into the issue, source registry, or run record.
7. Records and retention
Each issue is created with a process.json record containing candidate ranks, deterministic bucket scores and reasons, selected URLs, model and token telemetry, post-processing results, and the exact system prompt used for composition. This is operational evidence, not a dump of private configuration: secret values and full fetched documents are not included.
Daily collection owns source-fetch outcomes and a bounded private source-health history. Its Actions summary reports source failures and sources silent across four observations; collection artifacts survive for seven days, including failed attempts. Composition reads retained captures and does not invent source observations. A successful Cloudflare build check establishes deployment status, not a fresh browser visit or factual verification of the article.
Separate attempt journals preserve acquired inputs, the complete generation request including prior context, returned responses, retry errors, and terminal outcomes. Same-date attempts do not overwrite each other. The workflow retains these journals as artifacts for seven days; they are not published on the site. A completed response can be replayed without another model request when its saved inputs and request still match exactly. Local retention preserves recent evidence and never assumes an unterminated journal belongs to a dead process. Historical backfills do not rewrite current source-health observations.
The newest four issues keep the full record so current behavior can be inspected and tuned. As a fifth issue arrives, older records are compacted to run counts and telemetry; candidate tables, per-source diagnostics, selected URL lists, and prompt snapshots are removed. The published issue remains in the archive, while the current source roster and this page describe the live system.