Methodology
How a department’s faculty list becomes publication records, citation data, and peer comparisons.
- Codex / research team
- Find sources, prepare imports, inspect evidence, and review corrections.
- Gemini
- Read CVs, extract answers, judge uncertain matches, and investigate records.
- Program
- Check evidence, query sources, connect records, and calculate results.
Find people and CVs
The roster establishes who belongs to the department, including people whose CVs cannot be found.
Find the department
- Your request“Add Dartmouth Government”
- Codex / research teamFind the official faculty list
Check the department website and identify the relevant faculty roster.
- Faculty list found?
Continue with a CV search for each person.
Stop this department import until a usable roster is available.
For each academic on the roster
- Codex / research teamFollow faculty-page CV links
Check personal websites. If needed, search more widely online.
- CV found?
Verify the person, document completeness, and date.
Avoid duplicate files and attach it to the correct academic.
Continue to extractionThe academic stays in the database. Missing data does not mean zero publications. Research records candidate Scholar links and university pages.
Check the university page’s links, the profile’s name and affiliation, and verified email where required. Import supported profiles and publications; uncertain matches remain in review. Reuse cached pages before spending more searches.
Scholar recovery implemented; independent OpenAlex discovery remains separateWhat is automatic today? Department research and CV discovery are coordinated by Codex and the research team. Document processing begins with imported CVs. A separate bounded recovery command can verify saved Scholar candidates without a CV, fetch their data, and merge their publications. It checks university-page links, or matching verified institutional email plus publication titles from university pages or a name-matched ORCID/OpenAlex record. These identity checks are performed by code, without Gemini. Partial publication imports are labelled.
Extract and enrich
Gemini proposes structured information. The program checks it against the evidence and keeps track of its source.
Read the CV
- ProgramRead the PDF into numbered lines
Keep the document, page locations, and original text.
- GeminiMake an outline and answer CV questions
Identify sections and cite the lines supporting each answer.
- GeminiExtract publications and custom fields
Read the sections selected in the extraction settings.
- ProgramCheck the source evidence
Validate line references, titles, years, DOIs, and constrained fields. Flag uncertain results for review.
If coverage is incomplete: Gemini checks missed lines and rereads sections with remaining gaps.
Find matching source records
- Program · CrossrefFind and validate DOIs
Match bibliographic records against the CV’s title, year, and authorship.
- Program · OpenAlexResolve the academic’s identity
Use ORCID, publication identifiers, names, and affiliation evidence. Check the works against the CV.
- Program · OpenAlexFetch works and publication data
Store the available metadata, citation counts, and annual citation histories.
- Program · Google Scholar · optionalVerify and fetch the profile
Check candidate profiles against CV titles and identity evidence. Retrieve profile metrics and articles within the configured limits.
If several profiles pass: Gemini may help choose among acceptable candidates. It cannot override the verification rules.
Reconcile and save
- ProgramMatch records across sources
Join clear matches using DOIs, titles, years, and compatibility rules.
- Gemini · when neededJudge ambiguous matches
Propose which remaining records describe the same work.
- ProgramCheck the proposed matches
Apply safeguards and preserve review decisions. Uncertain matches remain flagged.
- ProgramSave publications and academic links
Keep the source of each chosen field, the original records, and the academic’s publication-specific answers.
If a usable profile cannot be found: keep the CV publications and record the unavailable or unresolved enrichment. Processing can continue with the available sources.
Your custom questions
The institution’s extraction settings define the questions and the sections Gemini reads.
- Questions about the academic
- For example: PhD year, PhD institution, or current position. Answers are saved on the academic with their CV evidence.
- Questions about each publication
- For example: “Is it peer reviewed?” or “Is this person the first author?” Answers stay on the academic–publication link, since coauthors may have different answers.
- Missing or uncertain answers
- Gemini is instructed to use unknown when the CV does not say. The program checks the cited evidence; these checks do not independently prove every interpretation.
What kind of faculty?
A separate appointment-enrichment step reads the CV’s job history.
- Gemini copies appointment titles, organizations, dates, and evidence.
- The program checks those lines and classifies rank, faculty flags, and administrative roles.
- An editor’s override takes priority. Otherwise verified CV appointments take priority over the CV’s current-position answer.
Scholar affiliation provides a cross-check. Conflicts go to review; missing evidence leaves the classification unknown.
What each external source supplies
- Crossref
- DOIs and bibliographic details, including journal and year, used to identify CV publications.
- OpenAlex
- Work identifiers, authorships, citations, annual citation history, journals, topics, open-access information, and other available metadata.
- Google Scholar
- Verified profile citation totals, annual citations, h-index, i10-index, interests, article lists, and available article citation counts.
Connect and count
The program builds relationships between the records, then calculates results using the selected counting rules.
From people to their work
- DepartmentA faculty roster in one discipline
- AcademicsMembers of the department
Connected to their publications through authorship links and custom answers.
- PublicationsCanonical records with source evidence
Choose a lead so related outputs count once.
Use identifiers and established spellings, plus available OpenAlex journal metadata.
From departments to comparisons
- Peer selectionSelect departments in the same field
The peer group is the set of departments chosen for comparison.
- ProgramApply the selected counting rules
Choose academics, publication sources, years, work types, citation scope, and journal weights where applicable.
- ProgramCalculate totals, per-academic values, or hybrid scores
Use the aggregation selected for each eligible graph.
- ResultsDatabase pages, analytics, and exports
Show the underlying sources and available coverage.
What belongs to each record?
- Academic
- Name, department, identifiers, CV, faculty classification, custom CV answers, profile matches, publications, citations, and available career history.
- Publication
- Title, year, type, status, identifiers, venue, authorship links, citations, provider metadata, and the source of each chosen value.
- Publication group
- Related outputs for an academic, such as an original, preprint, reprint, or translation. Members can keep their own rows while the group counts once. Citation calculations deduplicate Scholar clusters and OpenAlex records.
- Journal
- Standard name, alternate spellings, identifiers, available journal metadata, publication and academic counts, citation statistics, and formula or manual weights. Weights use stored data; Gemini does not assign them.
- Department
- Faculty membership and composition, publication output, citations, journals, and career statistics derived from its academics. The database calls a department a cohort.
- Peer group
- A selection of departments within a field in the same database workspace. Currently this is a comparison selection, rather than a separately stored peer-group entity.
Which publications and citations count?
Authorship and publication groups
A CV or verified Scholar profile must support authorship. OpenAlex-only attribution is retained but does not establish a counted publication.
Related outputs count once with their group. A shared publication counts for each coauthor, and once within their department’s total. Works in progress and non-publication items remain outside publication counts.
Profile totals and filtered works
Profile totals use the verified Scholar profile’s whole-career total, with OpenAlex citations of counted CV works as the fallback where permitted. Publication filters and journal weights do not change that total.
Filtered-work citations are calculated work by work. Journal weights apply here. Scholar and OpenAlex are alternative sources for a citation number; their totals are not added together.
Annual citations
An academic’s bar chart shows citations received in each calendar year. It uses verified Scholar history, or stored OpenAlex histories of counted CV works as the fallback. The department page shows each academic’s chart in a grid.
Missing data and journal weights
Unknown values remain unknown, rather than becoming zero. Journal formulas use statistics across the visible departments in the selected field; manual weights override the formula for individual journals. These weights affect the selected works’ citation scores.
Check with AI
The broader AI audit is a separate, budgeted review of assembled records. It is not automatically run after every import.
An inexpensive sweep reads the academic’s record and lists leads.
Use Google Search to confirm or reject leads and investigate issues the first pass missed.
Record the problem, evidence, confidence, and suggested correction.
- Codex / research teamReview the findings and their evidence
Decide whether a data correction or a change to the program’s rules is supported.
- Correction supported?
Investigate further when the available evidence warrants it.
The audit produces a report; it does not automatically edit records. It checks for wrong people or profiles, incorrect CVs, misattributed or missing works, duplicates, incorrect merges, wrong identifiers, incorrect metadata, and items that are not publications.
How API use is controlled
Stored source records and extraction results are reused where possible. Processing saves progress so interrupted work can resume. Scholar lookups have configured call limits, and provider costs are recorded.
The separate AI audit has its own scope and budget. Its Google Search investigation uses Gemini’s search tool; Scholar retrieval uses SerpAPI. Reading this page, changing analytics filters, and exporting stored data do not launch new Gemini or SerpAPI searches.