# Retrieval evaluation scorecard

Blank on purpose. Replace every placeholder with questions from your own workload. Do not publish customer mail, file names, or message text.

Score live Gmail and Drive search, the current index, and the proposed curated index against the same questions. Add a category to the index only when this sheet shows a gain you can point at.

| Id | Question class (live, curated, or both) | Question (yours) | Expected evidence (yours) | Who may see it | Correct | Citation matches source | Latency | Tool failure | Permission held | Cost notes |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| Q1 |  |  |  |  |  |  |  |  |  |  |
| Q2 |  |  |  |  |  |  |  |  |  |  |
| Q3 |  |  |  |  |  |  |  |  |  |  |

Acceptance thresholds come from this sheet after you fill it. There is no universal corpus size in this template.
