EDITORIAL SYSTEM / AI-ASSISTED RESEARCH

A source-aware AI writing workflow: citations are not enough

A source-aware article is not the draft with the most links. It is the article in which every material claim inherits authority from a checked source, commentary stays commentary, inference is labeled, and a missing fact can still stop publication.

Saved snapshotJuly 15, 2026
Saved rows37,627
Source types4 indexed surfaces
PublicationOne article

Last verified: . Details may be stale after this date. Research products, model behavior, benchmarks, and source-access rules change quickly. Recheck the linked documentation and research before reusing the implementation details or benchmark results.

Composition of a saved-source snapshot with 35,094 X posts, 2,000 YouTube Watch Later items, 477 Arc bookmarks, and 56 YouTube playlist items
Figure 1. The saved-source pool used for this editorial run. It is a point-in-time documents export from July 15, 2026, not a measure of search demand. The live refresh was unavailable, so the snapshot may omit newer saves. Counts are firsthand measurements from the local source index.

A cited draft is still a draft

Current research tools make source discovery much better. OpenAI documents that ChatGPT deep research can work across the public web, uploaded files, connected apps, and user-selected sites, then return a report with citations or source links. That is a useful research surface. It is not a guarantee that every adjacent sentence is supported by the cited page.

OpenAI's own accuracy guidance tells users to verify important information, including quotes, data, technical details, and references. A 2026 preprint, Cited but Not Verified, evaluated 14 closed and open models by checking whether citation links worked, contained relevant material, and actually supported the claim. In that study, the strongest models had link validity above 94% and relevance above 80%, but factual accuracy across evaluated systems ranged from 39% to 77%.

Those numbers belong to that paper's models, prompts, retrieval depth, and evaluation method. They are not a universal accuracy rate for every cited answer. The durable lesson is narrower: a working, relevant link and a supported claim are three different checks.

A citation gives you somewhere to look. Verification decides whether the sentence survives.

THE FIRST SEPARATION

Use saved sources to find the question, not to settle it

João's source index is deliberately broad. X likes capture fast commentary and demonstrations. Arc bookmarks retain pages. YouTube queues hold talks and implementation walkthroughs. This run found recurring interest in research agents, citation integrity, agent skills, MCP authorization, tool-use evaluation, and MCP Apps interfaces.

That recurrence is a topic signal, not audience demand and not proof. A saved post from Perplexity framed research as two jobs: search wide enough to find the relevant set, then investigate deeply enough to support every claim. It led to Perplexity's WANDR benchmark. The post was useful discovery context; the research page and its disclosed method are the citable source.

The same boundary applies to product behavior. A practitioner post can reveal a sharp failure mode or useful vocabulary. Current documentation from the product owner decides what the product claims to do. Primary research or a dated firsthand test supports measured outcomes. Commentary can then help explain why those outcomes matter.

Five-level source authority ladder from official documentation through primary research, engineering analysis, saved community sources, and explicit inference
Figure 2. The authority ladder used in this workflow. It is an original editorial model, not a platform ranking rule. A source's usable authority depends on the claim: product owners document product behavior, while independent measurements test outcomes.

THE SOURCE MAP

Map claims before writing paragraphs

The source map is the most important artifact in the workflow. It forces the research agent and editor to agree on what is being claimed, what kind of source could support it, when the source was checked, and where its authority ends.

ClaimSource and typeCheckedLimitation
ChatGPT deep research lets a user choose source sets and returns a cited reportOpenAI Help Center
Current official documentation
2026-07-16Documents product behavior; does not prove that each citation supports each claim
Citation link validity and claim support can divergeCited but Not Verified
2026 primary research preprint
2026-07-16Results are scoped to 14 systems and the paper's evaluation design
This run searched a snapshot containing 37,627 saved-source rowsLocal Convex documents export
Firsthand measurement
2026-07-15Point-in-time export; the blocked live refresh may mean newer rows are absent
Discovery and authority should be separate editorial stagesThis workflow
Explicit design judgment
2026-07-16An editorial method, not a documented search or model requirement

This format prevents source laundering. A social post cannot quietly become evidence for a protocol requirement. A vendor benchmark cannot become a universal model ranking. A primary study cannot escape its tested dataset. A firsthand result cannot turn one run into a market claim.

WIDE, THEN DEEP

Retrieval coverage is part of factual quality

A claim can be well supported by the wrong slice of the literature. Source-aware writing therefore needs both coverage and claim-level verification.

AutoResearchBench tests scientific literature discovery as two separate tasks: finding a specific target paper through multi-step search and collecting an unknown set of qualifying papers. Its strongest evaluated agents reached 9.39% accuracy on the deep task and 9.31% intersection-over-union on the wide task. That is a deliberately difficult scientific benchmark, not a score for general web research. It demonstrates why "the agent searched" is not evidence that the right evidence set was found.

DeepResearch Bench makes another useful separation. It evaluates report quality with one framework and information retrieval plus citation accuracy with another. Perplexity's DRACO benchmark likewise grades factual accuracy, breadth and depth, presentation, and primary-source citation as distinct dimensions. The benchmark comes from a product vendor and uses an LLM-as-judge protocol, so its claims should be read with those disclosed limits.

The practical editorial rule is simple: generate multiple candidate searches, record what each source contributes, and name the missing evidence. Do not reward the agent for source count alone.

Six-stage source-aware editorial workflow ending in either publication or a return to research when material evidence gaps remain
Figure 3. The workflow used for this article, drawn from the project editorial gate and this run's execution. It is original process documentation, last checked July 16, 2026.

THE REAL CONTROL

Let an evidence gap stop the article

The failure state of most writing automations is not an obviously fabricated citation. It is a polished draft that makes the editor feel almost done. Once prose becomes coherent, narrowing a claim or publishing nothing feels like throwing away work.

A source-aware workflow makes gaps explicit before that pressure arrives:

  • Authority gap: only commentary supports a product or protocol claim.
  • Coverage gap: the research found examples but not the relevant set or counterexample.
  • Freshness gap: the source is authoritative but may predate a fast-moving product change.
  • Provenance gap: a useful assertion cannot be mapped to the exact source passage or firsthand observation.
  • Visual gap: a screenshot, chart, or diagram would appear evidentiary but cannot be reproduced or attributed responsibly.

PaperTrail, a CHI 2026 study, explored a claim-evidence interface for scholarly question answering. Its interface exposed supported assertions, unsupported claims, and omissions. In a study with 26 researchers, that visibility lowered trust compared with the baseline, but the added caution did not consistently change reliance behavior. The result is a useful warning: showing provenance is valuable, but it does not remove the cognitive cost of verification or guarantee that a person will act on the warning.

That is why the gate must be procedural. If a material claim still has an authority, coverage, freshness, or provenance gap, the workflow narrows it, researches again, or publishes nothing.

RUN RECORD

Measured, documented, and inferred

StatementStatusEvidence and limit
The saved-source snapshot contained 37,627 rows across four indexed surfacesMeasuredJuly 15 documents export; counts exclude any newer rows not present in that snapshot
The production homepage and blog returned HTTP 200 before publicationMeasuredLive Vercel responses checked July 16; availability can change after the check
ChatGPT deep research exposes citations and source controlsDocumentedCurrent OpenAI product documentation; other products have different controls
Separating discovery from authority improves this editorial workflowInferenceMechanism is auditable here, but this run does not compare publication accuracy against another workflow

A reusable checklist

  1. Define one reader question. Record the direct answer and check the current site for intent collisions.
  2. Search personal and public sources. Use saved material for vocabulary, examples, and leads; use the current web to refresh and challenge them.
  3. Classify provenance. Separate official documentation, primary research, firsthand testing, practitioner analysis, community commentary, and inference.
  4. Build the claim map. Record the source, type, checked date, and limitation before drafting material claims.
  5. Search wide, then verify deep. Look for missing entities and counterexamples, then open the exact sources that support the final answer.
  6. Draft within the evidence boundary. Attribute commentary, scope measurements, and label every inference that could otherwise read as fact.
  7. Verify the rendered page. Check source links, structured data, visual provenance, internal links, mobile layout, and a prominent last-verified warning.
  8. Keep a no-publish branch. If a material gap remains, return the research batch instead of converting schedule pressure into a weak article.

What would change this conclusion?

A system that reliably retrieves the relevant source set, maps each material claim to the exact supporting passage, detects contradictions and omissions, and demonstrates those results across realistic independent evaluations could automate more of this gate. The current evidence does not show that the review step can be removed.

For now, the right division of labor is narrower. Let research agents expand the search, organize candidates, draft inside a source map, and expose gaps. Let the publication process remain capable of saying no.

Primary references and research context

Continue reading