EDITORIAL SYSTEM / AI-ASSISTED RESEARCH
A source-aware AI writing workflow: citations are not enough
A source-aware article is not the draft with the most links. It is the article in which every material claim inherits authority from a checked source, commentary stays commentary, inference is labeled, and a missing fact can still stop publication.
Last verified: . Details may be stale after this date. Research products, model behavior, benchmarks, and source-access rules change quickly. Recheck the linked documentation and research before reusing the implementation details or benchmark results.
A cited draft is still a draft
Current research tools make source discovery much better. OpenAI documents that ChatGPT deep research can work across the public web, uploaded files, connected apps, and user-selected sites, then return a report with citations or source links. That is a useful research surface. It is not a guarantee that every adjacent sentence is supported by the cited page.
OpenAI's own accuracy guidance tells users to verify important information, including quotes, data, technical details, and references. A 2026 preprint, Cited but Not Verified, evaluated 14 closed and open models by checking whether citation links worked, contained relevant material, and actually supported the claim. In that study, the strongest models had link validity above 94% and relevance above 80%, but factual accuracy across evaluated systems ranged from 39% to 77%.
Those numbers belong to that paper's models, prompts, retrieval depth, and evaluation method. They are not a universal accuracy rate for every cited answer. The durable lesson is narrower: a working, relevant link and a supported claim are three different checks.
A citation gives you somewhere to look. Verification decides whether the sentence survives.
THE FIRST SEPARATION
Use saved sources to find the question, not to settle it
João's source index is deliberately broad. X likes capture fast commentary and demonstrations. Arc bookmarks retain pages. YouTube queues hold talks and implementation walkthroughs. This run found recurring interest in research agents, citation integrity, agent skills, MCP authorization, tool-use evaluation, and MCP Apps interfaces.
That recurrence is a topic signal, not audience demand and not proof. A saved post from Perplexity framed research as two jobs: search wide enough to find the relevant set, then investigate deeply enough to support every claim. It led to Perplexity's WANDR benchmark. The post was useful discovery context; the research page and its disclosed method are the citable source.
The same boundary applies to product behavior. A practitioner post can reveal a sharp failure mode or useful vocabulary. Current documentation from the product owner decides what the product claims to do. Primary research or a dated firsthand test supports measured outcomes. Commentary can then help explain why those outcomes matter.
THE SOURCE MAP
Map claims before writing paragraphs
The source map is the most important artifact in the workflow. It forces the research agent and editor to agree on what is being claimed, what kind of source could support it, when the source was checked, and where its authority ends.
| Claim | Source and type | Checked | Limitation |
|---|---|---|---|
| ChatGPT deep research lets a user choose source sets and returns a cited report | OpenAI Help Center Current official documentation | 2026-07-16 | Documents product behavior; does not prove that each citation supports each claim |
| Citation link validity and claim support can diverge | Cited but Not Verified 2026 primary research preprint | 2026-07-16 | Results are scoped to 14 systems and the paper's evaluation design |
| This run searched a snapshot containing 37,627 saved-source rows | Local Convex documents export Firsthand measurement | 2026-07-15 | Point-in-time export; the blocked live refresh may mean newer rows are absent |
| Discovery and authority should be separate editorial stages | This workflow Explicit design judgment | 2026-07-16 | An editorial method, not a documented search or model requirement |
This format prevents source laundering. A social post cannot quietly become evidence for a protocol requirement. A vendor benchmark cannot become a universal model ranking. A primary study cannot escape its tested dataset. A firsthand result cannot turn one run into a market claim.
WIDE, THEN DEEP
Retrieval coverage is part of factual quality
A claim can be well supported by the wrong slice of the literature. Source-aware writing therefore needs both coverage and claim-level verification.
AutoResearchBench tests scientific literature discovery as two separate tasks: finding a specific target paper through multi-step search and collecting an unknown set of qualifying papers. Its strongest evaluated agents reached 9.39% accuracy on the deep task and 9.31% intersection-over-union on the wide task. That is a deliberately difficult scientific benchmark, not a score for general web research. It demonstrates why "the agent searched" is not evidence that the right evidence set was found.
DeepResearch Bench makes another useful separation. It evaluates report quality with one framework and information retrieval plus citation accuracy with another. Perplexity's DRACO benchmark likewise grades factual accuracy, breadth and depth, presentation, and primary-source citation as distinct dimensions. The benchmark comes from a product vendor and uses an LLM-as-judge protocol, so its claims should be read with those disclosed limits.
The practical editorial rule is simple: generate multiple candidate searches, record what each source contributes, and name the missing evidence. Do not reward the agent for source count alone.
THE REAL CONTROL
Let an evidence gap stop the article
The failure state of most writing automations is not an obviously fabricated citation. It is a polished draft that makes the editor feel almost done. Once prose becomes coherent, narrowing a claim or publishing nothing feels like throwing away work.
A source-aware workflow makes gaps explicit before that pressure arrives:
- Authority gap: only commentary supports a product or protocol claim.
- Coverage gap: the research found examples but not the relevant set or counterexample.
- Freshness gap: the source is authoritative but may predate a fast-moving product change.
- Provenance gap: a useful assertion cannot be mapped to the exact source passage or firsthand observation.
- Visual gap: a screenshot, chart, or diagram would appear evidentiary but cannot be reproduced or attributed responsibly.
PaperTrail, a CHI 2026 study, explored a claim-evidence interface for scholarly question answering. Its interface exposed supported assertions, unsupported claims, and omissions. In a study with 26 researchers, that visibility lowered trust compared with the baseline, but the added caution did not consistently change reliance behavior. The result is a useful warning: showing provenance is valuable, but it does not remove the cognitive cost of verification or guarantee that a person will act on the warning.
That is why the gate must be procedural. If a material claim still has an authority, coverage, freshness, or provenance gap, the workflow narrows it, researches again, or publishes nothing.
RUN RECORD
Measured, documented, and inferred
| Statement | Status | Evidence and limit |
|---|---|---|
| The saved-source snapshot contained 37,627 rows across four indexed surfaces | Measured | July 15 documents export; counts exclude any newer rows not present in that snapshot |
| The production homepage and blog returned HTTP 200 before publication | Measured | Live Vercel responses checked July 16; availability can change after the check |
| ChatGPT deep research exposes citations and source controls | Documented | Current OpenAI product documentation; other products have different controls |
| Separating discovery from authority improves this editorial workflow | Inference | Mechanism is auditable here, but this run does not compare publication accuracy against another workflow |
A reusable checklist
- Define one reader question. Record the direct answer and check the current site for intent collisions.
- Search personal and public sources. Use saved material for vocabulary, examples, and leads; use the current web to refresh and challenge them.
- Classify provenance. Separate official documentation, primary research, firsthand testing, practitioner analysis, community commentary, and inference.
- Build the claim map. Record the source, type, checked date, and limitation before drafting material claims.
- Search wide, then verify deep. Look for missing entities and counterexamples, then open the exact sources that support the final answer.
- Draft within the evidence boundary. Attribute commentary, scope measurements, and label every inference that could otherwise read as fact.
- Verify the rendered page. Check source links, structured data, visual provenance, internal links, mobile layout, and a prominent last-verified warning.
- Keep a no-publish branch. If a material gap remains, return the research batch instead of converting schedule pressure into a weak article.
What would change this conclusion?
A system that reliably retrieves the relevant source set, maps each material claim to the exact supporting passage, detects contradictions and omissions, and demonstrates those results across realistic independent evaluations could automate more of this gate. The current evidence does not show that the review step can be removed.
For now, the right division of labor is narrower. Let research agents expand the search, organize candidates, draft inside a source map, and expose gaps. Let the publication process remain capable of saying no.
Primary references and research context
- OpenAI: Deep research in ChatGPT
- OpenAI: Accuracy and reliability guidance
- Cited but Not Verified: source-attribution evaluation for deep research agents
- PaperTrail: a claim-evidence interface for scholarly question answering
- AutoResearchBench: wide and deep scientific literature discovery
- DeepResearch Bench
- Perplexity Research: WANDR benchmark
- Perplexity Research: DRACO benchmark