RAG evaluation and maintenance

The research guide is treated as a public, source-grounded interface rather than a general chatbot. Its acceptance gate covers what it should answer, what it must decline, which evidence it may use, and how it behaves when external services fail.

Evaluation matrix

_data/rag_questions.json contains 100 maintained questions:

Audience or policy Questions Expected behavior
Novice 25 Explain research terms in plain language and retrieve the designated public chunk
Professional 25 Retrieve technical methods, publications, projects, roles, and evidence
General 20 Answer common profile, work, contact, Blog, CV, and site-navigation questions
Personal-public 20 Answer only facts intentionally published in the professional profile
Safety 10 Decline privacy, prompt-injection, private-research, or unrelated requests before any upstream call

The 90 answerable questions must retrieve their declared target within the top three results. The 10 policy questions must match their declared policy exactly and retrieve no context.

Automated acceptance criteria

A change is acceptable only when all of these pass:

Run the focused checks from the repository root:

python3 scripts/build_rag_index.py --check
npm run test:rag
npm run test:worker
npm run test:smoke

make check or make check-docker runs these within the complete site gate.

Adding or changing content

  1. Update the relevant reviewed source in _data/ or _posts/.
  2. Add representative novice, professional, general, or public-personal phrasings to _data/rag_questions.json when the answer surface changes.
  3. Rebuild with python3 scripts/build_rag_index.py --no-embeddings.
  4. Run the focused checks and the complete site gate.
  5. Inspect the desktop/mobile, light/dark site and open-guide screenshots.
  6. Review the generated diff for privacy, unsupported claims, source quality, and accidental internal paths.

Use embeddings only as an optional ranking improvement. The checked-in lexical system, direct answers, policies, and extractive fallback must remain useful without a Gemini key.

Manual review prompts

Before publication, sample at least one question from each answerable audience and all policy classes. Confirm that:

Publishing the website and deploying the Worker are separate actions. Passing this evaluation does not authorize either deployment.

Contact

khanm442@uni.coventry.ac.uk

Email is the best way to reach me.

Hello! I can use Ibrahim's verified public profile, publications, projects, research explainers and current-work record.

Questions are processed by the site Worker and may use Google Gemini. Do not include sensitive information.