Skip to content

Source Evaluation and Evidence Policy

Branch: development Status: Active IKaC policy Related: Source Packet Workflow · Repository-wide refactor governance · Main Evidence & Verification Audit


Repository-wide evaluation (effective immediately)

Intake type ≠ refactor scope. A faculty profile, grant PDF, publication DOI, or ORS annual report may contribute evidence to any knowledge layer.

Every source packet evaluation must:

Requirement Policy
Full impact assessment Review all layers listed in Repository-wide refactor governance
Three-way classification Evidence-backed update · candidate knowledge · rejected claim
Candidate preservation Insufficient evidence → preserve in section M / extracted/candidates.yaml; never silent discard
Promotion review (mandatory) Every preserved candidate triggers targeted evidence discovery → re-evaluate → promote what becomes evidenced, or defer with rationale (section N). See Targeted evidence discovery and the Promotion Discovery Rule.
Layer-specific rules Apply grant, publication, or output rules when that layer is implicated — not as a packet scope limit

Core distinction (do not collapse)

Every source packet maintains two separate concepts:

Concept Meaning
Captured source A file or URL snapshot saved in the packet with provenance (HTTP metadata, timestamp, SHA-256, capture_status).
Source used as evidence A captured (or uploaded) source that supports one or more specific claims in evidence_checklist.md and proposed corpus edits.

Capture does not imply use. All captured sources are preserved for audit regardless of whether they are used as evidence. Refactor reports must list both sets explicitly.


Capture status

Intake records a capture_status for each URL snapshot:

Status Meaning Examples
success Expected content retrieved and stored. HTTP 200 HTML faculty profile; PDF upload saved intact.
partial Something was stored, but content is incomplete, blocked, or not the intended document body. HTTP 403 publisher landing page (error page or paywall shell saved for audit); truncated response; non-HTML body saved with note.
failed No usable snapshot stored. Network error; blocked URL; oversize response; validation failure.

HTTP 403 is always partial, never success. A saved 403 body is provenance (what the server returned at capture time), not confirmation that the scholarly record was retrieved. Titles, DOIs, and authorship for blocked publisher pages must come from higher-confidence sources in the same packet (e.g. faculty profile, Google Scholar, ORCID) unless independently verified.


High-confidence scholarly and institutional sources

These may support corpus updates when captured in a packet and cited in evidence_checklist.md:

  • University faculty profile pages
  • University laboratory websites
  • University department websites
  • University research center websites
  • Grant proposals
  • Grant award notices
  • Published journal articles
  • DOI landing pages
  • Google Scholar profiles
  • ORCID profiles
  • Web of Science records
  • Scopus records
  • OpenAlex records
  • DePaul Scholars / Elsevier Pure (scholars.depaul.edu) public OA portal — harvest authorized in sources/inbox/2026-08-20-library-harvest-authorization/ (HTML scrape until a Pure API/export is available)
  • VIA Digital Commons theses and dissertations — OAI-PMH metadata and OA full-text harvest under the same Library authorization
  • ResearchGate profiles when clearly associated with the scholar
  • Professional society profiles
  • Official project websites
  • Faculty-maintained research websites

Appropriate uses

High-confidence sources may support claims about:

  • People (identity, affiliation, contact when public)
  • Publications (titles, authorship, venues, years, DOIs when listed)
  • Collaborations (coauthorship, joint projects when explicitly stated)
  • Memberships (societies, consortia, editorial boards)
  • Service roles (advising director, journal editor, committee service) — on resource pages, only when tied to facilities or hardware (see below)
  • Expertise areas and research themes
  • Professional activities
  • Organizational participation

Scholarly index and profile systems

Google Scholar, ORCID, Web of Science, Scopus, OpenAlex, DePaul Scholars (Pure), and ResearchGate are valuable evidence sources for:

  • Publication lists and bibliographic metadata
  • Coauthor networks and collaboration patterns
  • Research themes and expertise areas
  • Professional identifiers (ORCID iD)
  • Cross-checking titles and DOIs when publisher pages are blocked

They are insufficient by themselves for:

  • Facility ownership or laboratory affiliation
  • Equipment ownership or specific instrument assignment
  • Facility access levels or booking policy
  • Operational status (whether a listed resource is staffed, maintained, and actually usable)
  • Organizational authority or administrative responsibility
  • Operational control of shared resources

Those relationships require stronger evidence from institutional or project sources (official lab pages, department facilities pages, grant award documents, explicit facility acknowledgments in publications or thesis methods/data chapters, access policies, or a documented curator operational assessment recorded in a source packet).

Thesis and dissertation full text, when lawfully captured, may support computing and data-handling practice claims and resource-use links when the methods, acknowledgments, or data chapter explicitly names a facility, instrument, storage method, or compute platform. A Scholars profile or a thesis author affiliation is not that evidence.

Named use vs corpus gaps (do not collapse). USB sticks, sneakernet, and similar transfer methods are rarely written down; zero lexicon hits is expected, not proof that removable media are unused. What is notable in a harvest cohort is missing research-practice evidence: no git/repository, no data-management tools (only analysis packages), no account of working or long-term storage or transfer, no data policy, no file names or locations, no supplemental methods/data appendix. Record those absences as cohort-level candidates (packet section M / unpublished reports) with the harvest scope (college, OA VIA set, date). Do not promote silence onto a person page or lab page as “this group has no DMP” or “this lab does not use git.”

Software and tools, including isolated use. Quoted analysis, scientific software, and everyday data infrastructure (Excel, Google Sheets/Docs, generic spreadsheets, CSV files, email attachments, paper surveys, codebooks) should be preserved even when the thesis names no lab, grant, or collaborator. An unconnected tool, person, course, or facility is still map evidence; ORS has found those isolates useful. When promoted, keep ordinary IKaC fields: access (red/yellow/green), verification/confidence, and traceable citations with excerpts. Do not drop a record because it does not create a collaboration edge. Gaps in the inventory belong in reports and Library/ORS notes, with harvest scope stated; they are not unsourced operational-status claims on a person or lab page.

Abstracts are a weak screen for this. CSH+CDM metadata found 35 practice candidates; full text found 257. LAS and Education abstracts often describe interviews and surveys without naming SPSS, Qualtrics, NVivo, or Excel. Treat “little LAS/Education computing in metadata” as methods-chapter not harvested, not as proof those colleges do not use software or data tools.

Computing questions may drive harvest; they do not shrink promotion. A computing-practice mine, unpublished draft, or ORS question is a valid intake frame. When a thesis is promoted into the corpus, existing IKaC still applies: capture, cite, and make visible the evidence; update every implicated layer; keep derivative views consistent. Shared instruments and computing sites can be collaboration evidence in the same sense as CDM-001 Idea Realization Lab. That is the ordinary full impact assessment, not a new ban on compute-focused work.

Faculty-maintained research websites sit between institutional pages and third-party indexes: strong for group research themes, news, and links; weaker alone for formal facility, access, or operational-status claims unless they repeat official institutional statements.

Licensed harvest vs promotion (do not collapse)

ORS-authorized Library ingest is licensed harvest into an archive, then curated promotion. It is not automatic corpus publication, and it is not the rejected v2 pattern of unattended HTML scraping as the map product.

Authorization is the source packet sources/inbox/2026-08-20-library-harvest-authorization/ (memo dated 2026-08-20). That memo records:

  • 12 August 2026 (ORS): Daniela Stan Raicu and Lauren Miller authorized expanding the Resource Map, collaborating with the Library and its databases, and assessing computing resources as available, needed, and actually used.
  • 19 August 2026 (Library): Kelly Hallisy, Ashley McMullin, and Kirsten Yehl authorized scraping scholars.depaul.edu and VIA theses/dissertations for the expanded map. The Library noted those corpora are incomplete and asked for help improving them.

The memo does not name OpenAlex, Scopus, Elsevier platform terms, or “PDF.” Harvest of Scholars HTML and VIA thesis/dissertation content proceeds on that Library authorization (curator reading: “scraping … VIA for thesis and dissertations” includes content, not metadata-only). OpenAlex (public API) and Scopus (institutional key, if issued) stay separately scoped.

Campus VPN (2026-08-21): licensed journal/proceedings full text for promotion of ranked candidates may be fetched while DePaul VPN is connected (scripts/check_campus_vpn.py). Scholars/VIA OA harvest does not require VPN. Do not bulk-download licensed corpora. If the VPN tunnel drops, stop publisher fetches and alert the curator to restart the client (see .cursor/rules/campus-vpn.mdc).

Store Location May contain Becomes map evidence when
Harvest archive sources/harvests/ (optional index under data/derived/) OAI-PMH metadata; OA HTML/PDF from Scholars and VIA under the packet above; OpenAlex/Scopus when separately keyed Never by itself
Curated corpus data/*.yaml, resources/ Reviewed records with quoted evidence Source packet (or equivalent checklist) + promotion review

Rules:

  • Capture does not imply use. Harvested records are preserved for audit and mining. They do not auto-create PER-* pages, PUB-* rows, or resource-use edges.
  • Scholars.depaul.edu (Pure): the public portal is OA. The Pure Web Services API returns 401 without an institutional key, and the Library cannot currently provision API access. OA portal scrape is the approved Scholars ingest path until an API or export exists. Prefer structured JSON/HTML already on the portal over inventing field mappings. Scholars robots.txt (User-agent: *) sets Crawl-Delay: 5 and disallows only ?format=rss and ?export=xls. Responses have carried tdm-reservation: 0 (text/data-mining rights not reserved).
  • VIA theses and dissertations: use OAI-PMH (https://via.library.depaul.edu/do/oai/) for structured metadata because the Library authorized this harvest, not because robots.txt prefers it. Generic User-agent: * on VIA disallows /do/ (the OAI path) and does not disallow /cgi/viewcontent.cgi (thesis PDFs). OA full-text harvest is in-scope under the packet's curator reading of the 19 August authorization. Elsevier Digital Commons OAI dataPolicy text that asks for written approval before robot full-text harvest is unchanged; this project proceeds on Library authorization in the packet, not on a conclusion that Elsevier's terms do not apply. Embargoed or login-walled files stay out.
  • OA harvest vs licensed promotion. Scholars/VIA OA harvest should succeed from the public internet. Licensed publisher PDFs for promotion may use campus VPN (2026-08-21). Bot-protection (Cloudflare / AWS WAF) is still not a reason to rotate UA/IP or solve CAPTCHAs.
  • Polite harvest (checkable). Identifiable User-Agent with a contact address; honor Scholars Crawl-Delay: 5; no User-Agent or IP rotation and no CAPTCHA-solving services; back off on HTTP 429/403; stop on operator request. A real browser session is allowed only to complete a WAF interstitial on an otherwise OA URL. Do not use it to evade paywalls or logins.
  • Identity match, then stop. Harvest records may be aligned to people.yaml by ORCID, email, or reviewed name+unit. Unmatched authors remain candidates.
  • Operational status is a separate layer (data/resource_operational.yaml) from access and verification. Public operational claims need packet-cited evidence. Curator knowledge that a service is unstaffed may be recorded as internal_notes and must not be paraphrased onto a public page as an unsourced fact.
  • Library-facing purpose. The 19 August meeting asked for help improving incomplete Scholars/VIA/database coverage, especially pending a Data Services Librarian hire. Sharing extraction that helps the Library is a later collaboration; it is not a public map claim.

See publications-layer-plan.md (2026-08-20 addendum) and sources/harvests/README.md.


Targeted evidence discovery (promotion review)

Candidate preservation is the beginning of promotion review, not the end. When a source produces candidates (facilities, labs, centers, institutes, clinics, equipment collections, capabilities, external facilities, programs, courses, partners, people, or relationships), the refactor must actively go look for corroborating evidence rather than leaving the candidate dormant. This is the Promotion Discovery Rule.

Permitted targeted-discovery sources (all subject to the confidence rules above):

Source Strong for Not sufficient alone for
Institutional websites (college/department/center/facility/project pages) facility existence, leadership, location, access — (these are the institutional evidence)
Faculty profiles & lab pages lab existence/leadership, people, research themes facility ownership beyond what the page states
Grant-related pages (e.g. ORS award listings) award existence, PI, program facility ownership / access
ORCID, Google Scholar, ResearchGate, Web of Science, Scopus, OpenAlex, DePaul Scholars (Pure) publications, coauthor networks, expertise facility / equipment ownership, access, or operational status
VIA / Digital Commons thesis metadata title, author, program, year, abstract resource use or computing practice unless methods text is captured and quoted
DOI-linked publications citation, authorship, funding acknowledgments facility ownership unless explicitly named in facilities text
Official partner websites external-partner existence, sub-facility structure, agreements DePaul-specific relationship unless stated

Rules for promotion review:

  • Capture before you cite. Every newly discovered source must be saved as a packet capture with provenance (URL, date, capture_status) before it backs a corpus edit.
  • The evidence bar is unchanged. Promotion of facility / lab / equipment / access / ownership / operational-status claims still requires institutional or project evidence; targeted discovery raises the obligation to look, not the threshold to promote. Harvest hits are discovery leads until quoted into a packet.
  • External partners may have internal structure. A national lab or large partner can contain multiple sub-facilities reached through different agreements and contacts; record which sub-facility is implicated and the access pathway, and do not assume one agreement covers all of them.
  • Record outcomes in section N (sources reviewed/imported, promoted/deferred, relationships + reciprocal links added, remaining gaps). Unresolved candidates stay candidates with reason_not_promoted.

Grant and award evidence (when the source documents awards)

When a source packet includes award or grant evidence, apply Grant ingestion governance for grant-layer decisions. This does not exempt the packet from repository-wide assessment — the same source may document labs, people, publications, or equipment.

Unresolved knowledge is still knowledge. A reliable award listing may enter data/grants.yaml even when the PI is not in people.yaml or facility links are unknown.

Treat as three independent decisions:

Decision Question
Grant existence Did the source document this award?
PI identity Is the named PI resolved to people.yaml (resolved / unresolved / ambiguous)?
Resource links Is there explicit evidence for a registry facility edge (resolved / unresolved / not_supported)?

Do not discard evidence-backed awards solely because the PI is absent from people.yaml. Do not invent person_id, people records, or facility links. Add person_grant_links.yaml only when person_id is resolved; add grant_resource_links.yaml only when facility naming is resolved.

Bulk grant sources: see Grant ingestion filter audit for current script behavior.


Publication registry (DOI-first, when implicated)

When a source packet or refactor identifies a scholarly publication (regardless of primary intake type):

  1. Resolve a DOI via Crossref, ORCID, Google Scholar, or publisher metadata — verify authorship and title match the discovery source (faculty profile, lab page, etc.).
  2. Store in data/publications.yaml standard citation fields only:
  3. title, year, venue, doi, url (https://doi.org/…)
  4. abstract (from Crossref, OpenAlex, or publisher abstract when available)
  5. people links to people.yaml
  6. evidence with discovery source + evidence_url DOI
  7. Do not save full-text PDFs or publisher HTML in the packet for registry intake — the DOI record plus verified citation metadata is sufficient for the Resource Map. Exception: the full-text facility-evidence pass may capture full text for a narrow, gated set of candidates solely to extract facility-use quotes. Captured full text is audit evidence under the source packet (fulltext/ or uploads/) with capture_status + SHA-256, is not registry content, and is not committed to the public repo when the license forbids redistribution. Publisher HTTP 403 remains partial.
  8. Link by DOI in resource prose and YAML (evidence_url, hyperlinks to https://doi.org/… or registry PUB-### IDs).
  9. Publisher pages returning HTTP 403 are partial captures only; DOI verification replaces publisher URL as the canonical reference.
  10. Add person_publication_links.yaml entries for verified DePaul authorship; add publication_resource_links.yaml only when explicit facility/instrument evidence exists (see publication-intake-rules.md).

ORCID registry (people.yaml)

ORCID iDs strengthen person identity and cross-check publication authorship. They do not support facility, equipment, or access claims by themselves.

When enriching or refactoring faculty/research-eligible people:

  1. Prefer Crossref authorship on linked publications (orcid_source: crossref, orcid_verification_status: verified).
  2. ORCID public API (name + DePaul affiliation, or email when public) may yield probable matches (orcid_verification_status: probable) — human confirm when ambiguous.
  3. Store on people.yaml:
  4. orcid — normalized iD (0000-0002-0542-6495)
  5. orcid_urlhttps://orcid.org/{orcid}
  6. orcid_verification_statusverified | probable | needs_review | unresolved | not_applicable (staff-only roles)
  7. orcid_verified_at — ISO date of last enrichment or manual confirm
  8. orcid_sourcecrossref | orcid_api | manual | pending
  9. orcid_candidates — list when multiple Crossref matches (needs_review)
  10. Batch enrichment: python3 scripts/enrich_scholarly_identifiers.py (report) or --apply (write YAML). Review docs/reports/scholarly-identifiers-review.md.
  11. Do not infer lab affiliation, instrument use, or resource access from ORCID alone.

Evidence citation targets (forward-only)

evidence_file / evidence_url on knowledge edges and nested associations must cite a primary source, not a derived knowledge object.

Allowed in evidence_file Forbidden as sole citation (terminal)
sources/inbox/… packet uploads and captures Uncaptured https://… / http://… web pages (except DOI records)
sources/harvests/… licensed harvest artifacts (see note) docs/**/*.md
sources/vault/…, sources/catalog/…, sources/extractions/… data/**/*.yaml / data/**/*.yml
DOI URLs (https://doi.org/…) resources/**/*.md or another link-registry YAML

A sources/harvests/… path may be a citation target only when a packet's evidence_checklist.md quotes the artifact. Harvest files are never sufficient by themselves to promote a YAML row, person page, or resource-use edge. The citation validator treats sources/ as primary; the packet quote is the promotion bar, not the path prefix.

Archive-first web evidence: Before a non-DOI web page supports a new or changed knowledge record, save a successful snapshot (or a partial capture that contains the actual source content) in a source packet and ingest it into the source catalog/vault. An HTTP error page may be retained for availability audit, but it cannot support the claim. Set evidence_file to that durable sources/… artifact and retain the current public page in evidence_url. DOI records remain an exception because the DOI and stored citation metadata are the durable identifier. Run python3 scripts/archive_evidence_urls.py --capture --rewrite --require-all to backfill cited pages. The forward validator blocks new uncaptured web evidence, and strict mode requires every non-DOI evidence_file to resolve to a present local artifact.

Forward-only rule: New or changed edges must not introduce terminal citations. Legacy terminals may remain until backfilled; they are reported by scripts/validate_evidence_citations.py and blocked in --mode forward when newly added or when an evidence pointer is changed to a terminal path. The scanner covers person/grant/publication/output link registries, people/grants/publications/ outputs records, and resource_ownership.yaml.

CSV/PDF under data/: Treated as export artifacts (allowed but not preferred). Prefer vaulting under sources/inbox/ and citing that path. Phase 2 vaulted the Internal Grants Summary CSV and NIH STRONG BRIDGE Facilities PDF into existing inbox packets; new edges should not reintroduce bare data/*.csv / data/*.pdf citations when a packet path exists. Phase 3 rewrites live https:// citations to packet captures only when a capture manifest / news capture / evidence-list sibling already provides a 1:1 map (DOI, catalog, and faculty profile URLs stay as https).

Validate locally / in review:

python3 scripts/validate_evidence_citations.py --mode forward   # default for PRs
python3 scripts/validate_evidence_citations.py --mode report    # full corpus inventory
python3 scripts/validate_evidence_citations.py --mode strict    # fail terminal or uncaptured web evidence

Repository audits under docs/reports/ (e.g. hypertext-reciprocity-audit.md) guide surfacing and linking work. Full rules: Audit-driven refactor policy.

  • Audit findings are not evidence.
  • Implement only changes supported by sources already in the repository or by a source packet.
  • Prefer surfacing YAML/registry evidence over inventing relationships to close audit gaps.

Service and professional activity on resource pages

Do not add service, editorial, or society-membership blocks to facility/lab/infrastructure resource pages unless the role directly involves facilities or hardware (e.g. director of a named core facility, instrument manager, machine-shop oversight).

Advising directorship, journal editorship, consortium membership, and similar roles:

  • May appear in people.yaml notes or person-oriented records when evidenced.
  • Must not clutter lab resource pages (e.g. CSH-016) unless facility-relevant.

Conservative claim categories

Remain conservative — require institutional or project-level evidence — for:

  • Facility ownership
  • Equipment ownership
  • Facility access
  • Operational status (staffed / maintained / actually usable as a service)
  • Organizational authority
  • Administrative responsibility
  • Laboratory affiliation (beyond what an official profile explicitly states)
  • Operational control

When evidence is indirect (method lists, coauthorship alone, shared building), use supporting language, needs_review, graph exclusion flags, or operational unknown/impaired with packet evidence — do not upgrade to definitive facility, access, or operational claims. Do not put unsourced curator judgments on public pages; use internal_notes.


Lower-confidence sources

Use with caution; generally not primary evidence unless corroborated:

Tier Examples Typical use
Lower confidence News articles, press releases, blogs, marketing materials Context, discovery leads, links to primary sources — corroborate before corpus edits
Very low confidence Social media, discussion forums, anonymous websites Do not use as primary evidence; may inform candidate links for separate intake

Refactor report requirements

Every refactor must produce refactor_output.md (or equivalent) with section A listing:

  1. Captured sources — every file under the packet (uploads/, source_snapshot/, source_snapshot/captured/) with capture_status where applicable.
  2. Sources used as evidence — which captures support which claims (map to evidence_checklist.md).
  3. Sources reviewed but not used — captures opened or listed but not cited in proposed edits.
  4. Reason not used — for each unused capture (e.g. "Google Scholar captured for audit; faculty profile used as primary evidence for publications"; "HTTP 403 publisher page — title taken from faculty profile instead").

Sections B–J, L (impact assessment), M (candidate registry), and N (promotion review summary) of the refactor output remain as defined in Source Packet Workflow and Repository-wide refactor governance. Section N is required whenever the packet produces candidates: it records the promotion-review outcome (sources reviewed/imported, candidates promoted/deferred, relationships and reciprocal links added, remaining gaps) or an explicit deferral with justification.


Curator-added provenance

Maintainer-added sources (outside the intake web form) are curator_added, not anonymous repository files. Each must record curator identity (added_by), date (added_at), and rationale (added_rationale) in packet metadata — see Source Packet Workflow § Curator-added sources.

Form-submitted sources are intake_form (intake_form_source in metadata.yaml). Captured source ≠ source used as evidence applies to both channels.


Hyperlinks extracted from URL snapshots (candidate_links.md) are not sources until a curator selects them for separate capture. After capture, each becomes a captured source subject to this policy — still distinct from "used as evidence" until the checklist and refactor cite it.


Lessons encoded (Kyle Grice packet, 2026-06-13)

  • A faculty profile URL intake can produce many candidate links; curators capture substantive links separately with full provenance.
  • Publisher HTTP 403 responses are partial captures — preserve the body and hash, do not treat as full article evidence.
  • Google Scholar may be captured and used for publications, collaborations, and expertise when associated with the scholar — but not as sole evidence for facility or equipment claims.
  • The official DePaul faculty profile remains the anchor for lab location and public affiliation; personal research sites supplement themes and contacts.
  • Refactors must document why a captured source was or was not used — silence is not acceptable audit practice.
  • Publications discovered on faculty profiles should be registered by verified DOI with citation metadata and abstract — not by publisher URL alone; do not store full text in packets.
  • Service roles (advising, editorship, IONiC, etc.) were removed from CSH-016 because they were not related to a documented resource or facility.