AI tool reviews should be useful, reproducible and honest about what was actually tested. This page explains how AI News & Updates plans, runs, scores, checks and corrects its reviews.
Methodology version 1.0. Last updated 31 July 2026.
Download the review rubric
The editable workbook contains the protocol, test cases, raw observation log, scoring weights, formula-driven summary and source/disclosure log. The PDF is a compact reference edition.
Our editorial principles
- Test before scoring. A product receives no score without a dated observation and supporting evidence.
- Keep tasks comparable. Products in the same comparison receive the same core prompt, eligibility rules, trial count, time limit and scoring weights. We record differences that make perfect parity impossible.
- Separate evidence from inference. We label hands-on observations, vendor documentation, independent evidence, editorial inference and unverified claims differently.
- Preserve failure. Outages, usage caps, broken citations, lost state and failed exports remain in the record. We do not silently retry until a tool looks good.
- Check material claims. A named human reviewer verifies source identity, relevance and support before publication.
- Disclose influence. Free access, affiliate relationships, sponsorship, vendor contact and correction requests are logged. None may alter the protocol or score.
- Correct visibly. Material corrections are dated, explained and reflected in the conclusion when necessary.
What we test
Every review begins with a written use case and explicit success criteria. We choose tasks that reflect what the intended reader is likely to do, not a theatrical prompt designed only to make the product fail. Where possible, we run at least three independent trials. Each trial starts in a clean project, collection or conversation, and we record the date, plan, model or mode, material settings, time spent, cost and any usage limit encountered.
For tools that retrieve or summarise external evidence, we sample material claims and trace them to the cited source. A citation must identify a real item, link to the correct item and actually support the nearby claim. A valid-looking DOI attached to the wrong conclusion is a citation failure. Access barriers are recorded separately from mismatched or fabricated references.
We also test recovery. We correct one mistaken constraint, trigger or encounter one failed action, and examine whether the tool preserves understandable state. We then export the work and check whether citations, notes and essential context survive outside the product.
The seven-part scorecard
| Criterion | Weight | What it measures |
|---|---|---|
| Task completion | 25% | Whether the tool completes the stated task accurately within the constraints. |
| Evidence traceability | 20% | Whether material claims are readily traceable to matching, accessible sources. |
| Reliability and recovery | 15% | Consistency, state preservation and clarity when something goes wrong. |
| Privacy and data handling | 10% | Documented retention terms, controls and handling of uploaded material. |
| Security and control | 10% | Account, sharing and administrative controls appropriate to the tested audience. |
| Cost value | 10% | Useful, verifiable outcomes at a transparent and proportionate total cost. |
| Maintenance | 10% | Export quality, change communication and ongoing correction or upkeep burden. |
Each criterion is scored from 0 to 10. Broad anchors are: 0–2 for material failure; 3–4 for partial completion needing substantial rescue; 5–6 for core completion with meaningful limitations; 7–8 for strong, verifiable performance with manageable limits; and 9–10 for excellent, repeatable performance with clear controls and evidence. Blank means untested, not zero.
The weighted score is the sum of each completed criterion score multiplied by its published weight. We do not publish a ranked winner until every included product has equivalent completed trials. The spreadsheet keeps raw observations separate from calculated summaries so readers can see where a score came from.
Our five benchmark test cases
- T01 — Search coverage: collect candidate results and check relevance, traceability, identifiers and duplicates.
- T02 — Evidence traceability: trace five material claims or summaries to their cited sources.
- T03 — Screening consistency: apply disclosed eligibility rules to ten borderline records and record reasons.
- T04 — Synthesis and disagreement: ask the tool to separate supporting, null and conflicting evidence.
- T05 — Recovery and export: correct a constraint, recover from a failed action and export the work.
Current research-assistant benchmark
Our current protocol compares ResearchRabbit, Elicit, Scite, Consensus and Undermind using this question: For university students, does retrieval practice improve delayed retention compared with rereading?
Eligible records must directly compare retrieval practice, practice testing or active recall with rereading or restudy in higher-education students and measure retention at least 24 hours later. Randomised or quasi-experimental comparisons are eligible. The standard prompt asks each product to identify study design, sample, delay, outcome and direction of effect; provide persistent identifiers or direct paper links; separate cited claims from synthesis; surface disagreement; and state what may have been missed.
No comparative result is published yet. The workbook is preloaded with the five products and five test cases, but all observations remain “Not Started” until a real, dated test is completed. This prevents a polished methodology page from being mistaken for a hands-on result.
Evidence labels
| Label | Meaning |
|---|---|
| Hands-on test | Observed directly in a dated test; export or screenshot preserved. |
| Vendor documentation | A supplier claim supported by a current URL and access date. |
| Third-party evidence | An independent source with author, publication and URL recorded. |
| Inference | A reasoned editorial conclusion clearly labelled as inference. |
| Unverified | A claim that could not be corroborated and is not presented as fact. |
Vendor contact, access and corrections
We may contact a vendor before or after publication to check factual details such as the official domain, current pricing, feature names or documentation. We welcome specific, sourced corrections. Vendors do not receive approval rights over our wording, test record, score or verdict.
If a vendor provides free or extended access, that is recorded in the source and disclosure log. We identify the plan actually tested and do not assume that controls listed for an enterprise plan exist on a free plan. If a test hits a usage limit, that limitation remains part of the observation.
Affiliate relationships, sponsorship and paid placement are disclosed near the relevant content. Payment cannot buy a dofollow link, a favourable review, a higher score, removal of a material limitation or access to editorial approval. We prefer useful, earned references and decline low-quality directory or link-scheme opportunities.
Publication and correction policy
Before publication, a named human reviewer checks the recorded plan and mode, equivalent trial counts, weights, sampled citations, disclosures and wording of material conclusions. A ranking remains unpublished if comparable testing is incomplete. We also state when a conclusion is limited by account access, missing exports, language restrictions, time limits or an incomplete human gold set.
Material corrections include a date, the original affected wording or score, the corrected information, the supporting evidence and any impact on the verdict. Minor spelling or formatting changes may be fixed without a formal correction note when they do not alter meaning.
Questions, evidence or correction requests can be sent to hello@ainewsandupdates.com. For an example of a vendor-supplied factual correction that did not alter our editorial judgement, see our ResearchRabbit review.