Pilot AI for TOR review: define the task, verify sources and measure results
A document-review pilot for IT and procurement teams, with a sample prompt, evaluation criteria and a downloadable results worksheet.
Start with a TOR document approved for testing. Have a person create reference requirements before asking AI to extract the same information. Check completeness, page references and human correction time. This is a proposed method and illustrative scenario, not a customer benchmark or a claim of model accuracy.
Best for: IT, procurement and project teams evaluating AI-assisted document reviewSeparate coverage from correctness
Count unique reference requirements and verify extracted items individually.
Build a reference
A reviewer lists requirements before using AI
Match items
Count correct, unique matches only
Review critical items
Flag errors that affect acceptance
Calculate from your own pilot
Enter reviewer-verified counts. Count each requirement once; empty fields are not treated as zero.
Complete all 3 fields to see your results.
N/A means a denominator is missing. These metrics assess extraction against the reference set, not every aspect of quality, and do not replace human review of critical requirements.
Compare before deciding
1. Prepare documents and reference answers
Use documents approved for this purpose and remove confidential and personal information before uploading. Have a reviewer list requirements, delivery terms, service conditions and acceptance criteria with page references. For scans, evaluate OCR quality separately from AI reasoning.
2. A prompt to adapt
Extract requirements only from the attached document. Return: category | requirement | conditions/exceptions | page | short supporting quote. Do not add requirements from general knowledge. If missing, write ‘not found in the document’. Separate ambiguous items for human review. Do not decide compliance on behalf of the responsible reviewer.
3. Check errors before speed
Verify important items against the original. Check multi-page tables, units, quantities, dates and minimum/maximum wording. Classify omissions, incorrect extraction and incorrect citations. Define unacceptable errors before testing rather than adjusting the criteria to fit the result.
4. Evaluate on documents not used to tune the prompt
Separate prompt-development documents from evaluation documents and include the formats your team encounters. Record model name, version, date, prompt and OCR steps. Re-evaluate using the same criteria after changes. Success on one document does not establish support for every document.
5. Decide using accepted results
Compare total time and errors against the current process. Keep human review where errors persist and limit the scope accordingly. Unaccepted AI output should not feed procurement approvals automatically. The downloadable worksheet leaves actual results blank for your team; it contains no invented scores.
Quotation preparation checklist
- Document use is authorized for the selected service
- Reference requirements and a reviewer are assigned
- Unacceptable errors are defined before testing
- Development and evaluation sets are separated
- Version, date, time and costs are recorded
- A responsible person accepts the output before use
Frequently asked questions
Can AI replace reading the TOR?
Use it to assist retrieval and categorization. The responsible person should verify important terms against the original before deciding.
Does this article contain actual benchmarks?
No. It is a method and worksheet for running your own pilot. Actual results must come from your documents and process.
Sources used for verification
Vendor facts are separated from TechTouch guidance. Features, pricing, and commercial terms may change, so confirm the latest information at the time of purchase.