Liberty91

Quality of Information Check.

/quality-of-information-check

The quality-of-information-check skill grades evidence, so it makes it easier to base assessments on it. It works by giving it a link to a (security) article, and it finds the primary sources referenced in the body of the text, breaks that report into separate claims, and grades each claim with two words: how the source knows, and what backs the statement. As with the rest of this skills pack, it is free, open source, and works on any article URL with no API key. It works with Claude Code, Cursor, Codex or Windsurf.

Why grade the evidence first.

Structured analytic techniques (SAT) are the methods intelligence analysts use to keep their reasoning objective: write down every possible explanation, test each one against the evidence, check what you are assuming, and check the quality of your information before you rely on it. The skills put those techniques into the AI's workflow so it has to follow them. One of the big problems with relying on AI to do this though, is that it can 'choose' to trust a source, which makes the analysis look rigorous. But in fact it's not. Grading the evidence first, by fixed rules, with every step checkable, is what makes the analysis more reliable.

How it works.

When you ask how reliable a piece of security news is, the skill works through seven steps in a fixed order. It does not grade before the source is found, and it does not count confirmations after grading.

  1. Find the original source. A small script follows the links in the article to the report it is based on. Fifteen articles can all lead to one report. That is one source, not fifteen. If the article links to nothing, the script flags that, and the AI is not allowed to fill the gap from memory.
  2. Break the report into claims. A sentence like "a criminal group broke into a hospital through a flaw in its network equipment and stole 400 gigabytes of data" is four claims from four different places: the hospital's own notice, the investigators' finding, the investigators' judgement about who did it, and the criminals' own number.
  3. Check whether the wording changed between the original and the article. "Possibly linked to" often becomes "was behind" in a news article about that primary report. The claim is graded on the original wording, and the way the news outlet covers it is recorded.
  4. Count who else confirms the event. A confirmation can only be from a second organisation with its own evidence. Several articles about the same report do not count, and neither does one company quoting another.
  5. Grade each claim on two axes, each with a reasoning and the exact wording it's based on, copied word for word from the report.
  6. Apply fixed limits. A script checks every grade against the published rules and rejects any that break them.
  7. Only then reason. Hypothesis testing and assessments are based on the graded claims, and those can never be more confident than the weakest claim.

The grading.

Every claim is measured against two axes, written as two words. The first word is the access level. It answers: how does the source know this? The second word is the claim support. It answers: what backs this statement?

Access level: how does the source know this?

WordMeaning
directThe source saw it: in its own systems, in its own investigation, or because it is the organisation that was attacked
limitedThe source examined a piece of evidence, such as a malware sample, but did not see the attack itself
indirectThe source does not say how it knows, or is repeating what others published
untracedNo original source could be found
adversaryThe attackers said it themselves

Claim support: what backs this statement?

WordMeaning
establishedA second organisation confirms it from its own evidence, or you saw it in your own systems
firmOne source, which saw it directly, and nothing contradicts it
tentativeOne source, and the statement is second-hand or is a conclusion the source reached
disputedSomething contradicts it, or the source itself says it has low confidence
unverifiedNothing to judge it by

Two claims from one report can get different grades.

In the hospital example, what the investigators saw in the systems is "direct, firm". The attribution to a threat actor is "direct, tentative". Both claims come from the same source. The difference is between what a source saw, and what a source concluded. Nobody saw the threat actor actually do it, but the investigators compared the tools and methods with what they know about that group and reached a conclusion. Another investigator with the same evidence could reach a different one, so one source's conclusion is "tentative" at best. To reach "established", any claim needs a second organisation that confirms it from its own evidence. And what the attackers say about how much they stole is "adversary, unverified" until someone else confirms it.

Why are we not using the Admiralty scale?

Analysts familiar with the Admiralty scale might wonder why we're not using that here. That's because it works really well for Human Intelligence (HUMINT), but not so well for Threat Intelligence Reporting. The grading we're proposing is a better fit.

In the Admiralty Scale, the letter is the source's track record: how often that source has been right before. Rating a track record takes a history of past claims and outcomes. Nobody keeps that history for hundreds of vendors and news outlets, and without that track record, "usually reliable" means it becomes an impression of a company's reputation, which is a typical analytical bias we're trying to prevent here.

Instead the grading system records what can be read from the report: how the source knows. A report either says "we observed this in our telemetry" or it does not. Two people reading the same report will reach the same answer.

Scripts and AI-checks.

Following the links to the original report, recording who published what, and checking every grade against the limits are done by scripts. The provenance tracing is a plain script, standard library only, and gives the same answer every time.

Deciding what kind of claim a statement is, judging whether an article changed the original wording, and writing the reason for each grade are judgements the AI-model makes. This is tracked and reported in the output, so you can review and check. Every grade is based on evidence, and a script checks that the sentence is really in the report. The grading rules and reasoning are published in the repository.

How to fit it in your workflow.

By default, the skills that test hypotheses and write assessments, /ach, /threat-assessment and /writing-assessments, grade their evidence. You can invoke them with a link, and the check will run before the analysis. You can skip it. The result is then labelled "ungraded" and cannot claim more than moderate confidence. If an indicator later shows up in your own logs, through /lookup-sentinel for example, or your own endpoint tooling, that observation is logged in the claim table as "direct, established": you saw it yourself, so nobody else needs to confirm it. The skill runs from Claude Code, Cursor, Codex or Windsurf, on any article URL, with no API key.

Frequently Asked Questions.

Want to see this in your own organisation?

Request a demo, or start for free right away, and you will be looking at actionable, real-time, relevant threat intelligence built around your own organisation instead of generic reporting.