Methodology

How a department’s faculty list becomes publication records, citation data, and peer comparisons.

Codex / research team
Find sources, prepare imports, inspect evidence, and review corrections.
Gemini
Read CVs, extract answers, judge uncertain matches, and investigate records.
Program
Check evidence, query sources, connect records, and calculate results.
1

Find people and CVs

The roster establishes who belongs to the department, including people whose CVs cannot be found.

From a request to an imported CV

Find the department

  1. Your request“Add Dartmouth Government”
  2. Codex / research teamFind the official faculty list

    Check the department website and identify the relevant faculty roster.

  3. Faculty list found?
Yes
Codex / research teamPrepare the roster
ProgramCreate or update the department and academics

Continue with a CV search for each person.

No
Codex / research teamRecord the search outcome

Stop this department import until a usable roster is available.

What is automatic today? Department research and CV discovery are coordinated by Codex and the research team. Document processing begins with imported CVs. A separate bounded recovery command can verify saved Scholar candidates without a CV, fetch their data, and merge their publications. It checks university-page links, or matching verified institutional email plus publication titles from university pages or a name-matched ORCID/OpenAlex record. These identity checks are performed by code, without Gemini. Partial publication imports are labelled.

2

Extract and enrich

Gemini proposes structured information. The program checks it against the evidence and keeps track of its source.

Read down each phase, then continue to the next: read the CV, find source records, and reconcile them.

Read the CV

  1. ProgramRead the PDF into numbered lines

    Keep the document, page locations, and original text.

  2. GeminiMake an outline and answer CV questions

    Identify sections and cite the lines supporting each answer.

  3. GeminiExtract publications and custom fields

    Read the sections selected in the extraction settings.

  4. ProgramCheck the source evidence

    Validate line references, titles, years, DOIs, and constrained fields. Flag uncertain results for review.

If coverage is incomplete: Gemini checks missed lines and rereads sections with remaining gaps.

Find matching source records

  1. Program · CrossrefFind and validate DOIs

    Match bibliographic records against the CV’s title, year, and authorship.

  2. Program · OpenAlexResolve the academic’s identity

    Use ORCID, publication identifiers, names, and affiliation evidence. Check the works against the CV.

  3. Program · OpenAlexFetch works and publication data

    Store the available metadata, citation counts, and annual citation histories.

  4. Program · Google Scholar · optionalVerify and fetch the profile

    Check candidate profiles against CV titles and identity evidence. Retrieve profile metrics and articles within the configured limits.

If several profiles pass: Gemini may help choose among acceptable candidates. It cannot override the verification rules.

Reconcile and save

  1. ProgramMatch records across sources

    Join clear matches using DOIs, titles, years, and compatibility rules.

  2. Gemini · when neededJudge ambiguous matches

    Propose which remaining records describe the same work.

  3. ProgramCheck the proposed matches

    Apply safeguards and preserve review decisions. Uncertain matches remain flagged.

  4. ProgramSave publications and academic links

    Keep the source of each chosen field, the original records, and the academic’s publication-specific answers.

If a usable profile cannot be found: keep the CV publications and record the unavailable or unresolved enrichment. Processing can continue with the available sources.

Your custom questions

The institution’s extraction settings define the questions and the sections Gemini reads.

Questions about the academic
For example: PhD year, PhD institution, or current position. Answers are saved on the academic with their CV evidence.
Questions about each publication
For example: “Is it peer reviewed?” or “Is this person the first author?” Answers stay on the academic–publication link, since coauthors may have different answers.
Missing or uncertain answers
Gemini is instructed to use unknown when the CV does not say. The program checks the cited evidence; these checks do not independently prove every interpretation.

What kind of faculty?

A separate appointment-enrichment step reads the CV’s job history.

  1. Gemini copies appointment titles, organizations, dates, and evidence.
  2. The program checks those lines and classifies rank, faculty flags, and administrative roles.
  3. An editor’s override takes priority. Otherwise verified CV appointments take priority over the CV’s current-position answer.

Scholar affiliation provides a cross-check. Conflicts go to review; missing evidence leaves the classification unknown.

What each external source supplies
Crossref
DOIs and bibliographic details, including journal and year, used to identify CV publications.
OpenAlex
Work identifiers, authorships, citations, annual citation history, journals, topics, open-access information, and other available metadata.
Google Scholar
Verified profile citation totals, annual citations, h-index, i10-index, interests, article lists, and available article citation counts.
3

Connect and count

The program builds relationships between the records, then calculates results using the selected counting rules.

How the database entities connect

From people to their work

  1. DepartmentA faculty roster in one discipline
  2. AcademicsMembers of the department

    Connected to their publications through authorship links and custom answers.

  3. PublicationsCanonical records with source evidence
Related outputs
Program · publication groupsConnect versions of a work

Choose a lead so related outputs count once.

Published in
Program · journalsResolve the journal

Use identifiers and established spellings, plus available OpenAlex journal metadata.

From departments to comparisons

  1. Peer selectionSelect departments in the same field

    The peer group is the set of departments chosen for comparison.

  2. ProgramApply the selected counting rules

    Choose academics, publication sources, years, work types, citation scope, and journal weights where applicable.

  3. ProgramCalculate totals, per-academic values, or hybrid scores

    Use the aggregation selected for each eligible graph.

  4. ResultsDatabase pages, analytics, and exports

    Show the underlying sources and available coverage.

What belongs to each record?

Academic
Name, department, identifiers, CV, faculty classification, custom CV answers, profile matches, publications, citations, and available career history.
Publication
Title, year, type, status, identifiers, venue, authorship links, citations, provider metadata, and the source of each chosen value.
Publication group
Related outputs for an academic, such as an original, preprint, reprint, or translation. Members can keep their own rows while the group counts once. Citation calculations deduplicate Scholar clusters and OpenAlex records.
Journal
Standard name, alternate spellings, identifiers, available journal metadata, publication and academic counts, citation statistics, and formula or manual weights. Weights use stored data; Gemini does not assign them.
Department
Faculty membership and composition, publication output, citations, journals, and career statistics derived from its academics. The database calls a department a cohort.
Peer group
A selection of departments within a field in the same database workspace. Currently this is a comparison selection, rather than a separately stored peer-group entity.

Which publications and citations count?

Authorship and publication groups

A CV or verified Scholar profile must support authorship. OpenAlex-only attribution is retained but does not establish a counted publication.

Related outputs count once with their group. A shared publication counts for each coauthor, and once within their department’s total. Works in progress and non-publication items remain outside publication counts.

Profile totals and filtered works

Profile totals use the verified Scholar profile’s whole-career total, with OpenAlex citations of counted CV works as the fallback where permitted. Publication filters and journal weights do not change that total.

Filtered-work citations are calculated work by work. Journal weights apply here. Scholar and OpenAlex are alternative sources for a citation number; their totals are not added together.

Annual citations

An academic’s bar chart shows citations received in each calendar year. It uses verified Scholar history, or stored OpenAlex histories of counted CV works as the fallback. The department page shows each academic’s chart in a grid.

Missing data and journal weights

Unknown values remain unknown, rather than becoming zero. Journal formulas use statistics across the visible departments in the selected field; manual weights override the formula for individual journals. These weights affect the selected works’ citation scores.

4

Check with AI

The broader AI audit is a separate, budgeted review of assembled records. It is not automatically run after every import.

Audit findings lead to reviewed corrections
Codex / research teamSelect the records and set an audit budget
Gemini · first passIdentify possible problems

An inexpensive sweep reads the academic’s record and lists leads.

Gemini · investigationCheck the whole record against the web

Use Google Search to confirm or reject leads and investigate issues the first pass missed.

Gemini · reportStructure the findings

Record the problem, evidence, confidence, and suggested correction.

  1. Codex / research teamReview the findings and their evidence

    Decide whether a data correction or a change to the program’s rules is supported.

  2. Correction supported?
Yes
Codex / editorApply the supported correction
ProgramRebuild affected results and run relevant checks
No or uncertain
Codex / research teamKeep the uncertainty documented

Investigate further when the available evidence warrants it.

The audit produces a report; it does not automatically edit records. It checks for wrong people or profiles, incorrect CVs, misattributed or missing works, duplicates, incorrect merges, wrong identifiers, incorrect metadata, and items that are not publications.

How API use is controlled

Stored source records and extraction results are reused where possible. Processing saves progress so interrupted work can resume. Scholar lookups have configured call limits, and provider costs are recorded.

The separate AI audit has its own scope and budget. Its Google Search investigation uses Gemini’s search tool; Scholar retrieval uses SerpAPI. Reading this page, changing analytics filters, and exporting stored data do not launch new Gemini or SerpAPI searches.