Checked 29 September 2026. From CT-AI Syllabus v2.0 (GA 17 Apr 2026) sections 3.3.1, 4.2.2, 5.1.5 and 6.1.5, and the Exam Structures & Rules Tables v1.19. Current as checked on 29 September 2026; re-check if ISTQB updates any of the cited documents (syllabus, sample exam, exam tables or FAQ).
CT-AI v2.0 has 43 learning objectives. 39 ask you to understand (K2). Only four ask you to apply (K3), and each of those is examined by one question worth two points. Together they are 8 of the 44 points from 4 of the 40 questions, and they are the only questions explicitly set at the apply level; the other 36 are K2 understanding questions, which can still ask you to compare, classify or infer. This guide takes the four in syllabus order, using the syllabus's own definitions and examples. We do not reproduce exam questions; ISTQB's free Sample Exam A v2.2 is where to see how they are asked.
CT-AI Syllabus v2.0; ISTQB Exam Structures & Rules Tables v1.19 and Exam Structures and Rules v1.2 (6.1.3).
| K3 objective | Syllabus wording | Chapter (points) | Section |
|---|---|---|---|
| AI-3.3.1 | Calculate common ML functional performance metrics from a given set of confusion matrix data | 3. Machine Learning (8) | 3.3.1 |
| AI-4.2.2 | Implement red teaming for GenAI systems | 4. Testing AI-Based Systems (8) | 4.2.2 |
| AI-5.1.5 | Apply dataset constraint testing | 5. Input Data Testing (7) | 5.1.5 |
| AI-6.1.5 | Use metamorphic testing to derive test cases for a given scenario | 6. Model Testing (10) | 6.1.5 |
CT-AI Syllabus v2.0 learning objectives; each K3 objective has one question worth 2 points in the Exam Structures & Rules Tables v1.19.
1. Confusion-matrix metrics (AI-3.3.1)
The syllabus draws the confusion matrix with Predicted on the rows and Actual on the columns, and warns that it "may be presented differently (e.g., predicted and actual swapped)". Before you calculate anything, find which axis is which: a swapped matrix puts FP and FN in each other's cells, and precision and recall change places.
| Actual positive | Actual negative | |
|---|---|---|
| Predicted positive | True Positive (TP) | False Positive (FP) |
| Predicted negative | False Negative (FN) | True Negative (TN) |
Layout of Figure 2 in CT-AI Syllabus v2.0, section 3.3.1.
| Metric | Formula (syllabus) | What it answers |
|---|---|---|
| Accuracy | (TP + TN) / (TP + TN + FP + FN) × 100% | Percentage of all classifications that were correct |
| Precision | TP / (TP + FP) × 100% | Of the predicted positives, how many were right: how sure you can be of a positive prediction |
| Recall (sensitivity) | TP / (TP + FN) × 100% | Of the actual positives, how many were found: how confident you can be that no positives are missed |
| F1-score | 2 × (Precision × Recall) / (Precision + Recall) | Harmonic mean of precision and recall, "with values ranging from 0 to 100" |
CT-AI Syllabus v2.0, section 3.3.1. Because precision and recall are already percentages, v2.0's F1-score comes out on the 0-100 scale too. Many ML textbooks and libraries give F1 between 0 and 1; in this exam, follow the syllabus.
A worked example, the same numbers as the calculator below
| Step | Working | Result |
|---|---|---|
| Inputs | TP 40, FP 10, FN 20, TN 130; total 200 | |
| Accuracy | (40 + 130) / 200 × 100% | 85.0 |
| Precision | 40 / (40 + 10) × 100% | 80.0 |
| Recall | 40 / (40 + 20) × 100% | 66.7 |
| F1-score | 2 × 80.0 × 66.7 / (80.0 + 66.7), using unrounded values | 72.7 |
| Trap: plain average of precision and recall | (80 + 66.666…) / 2, using the unrounded recall | 73.3 (not the F1-score) |
| Trap: matrix read with axes swapped | FP and FN change places | precision 66.7, recall 80.0 |
Our example and arithmetic, using the syllabus formulas. Each result is computed from unrounded values and then rounded to one decimal place (recall is 66.666…, shown as 66.7). The numbers are invented for practice and are not from any ISTQB paper.
Confusion-matrix calculator, on the syllabus's 0–100 scale
Enter the four cells of a confusion matrix. The calculator applies the formulas exactly as CT-AI Syllabus v2.0 section 3.3.1 writes them, multiplied by 100%, and shows the arithmetic-mean mistake many candidates make with F1.
Formulas: ISTQB CT-AI Syllabus v2.0, 3.3.1. Rounded to one decimal place for display. In the exam, read which way round the matrix is drawn: the syllabus notes predicted and actual may be swapped.
Two habits pay off here. First, write the four cell values down before you touch a formula: a common error is misreading which axis is predicted and which is actual. Second, when precision and recall differ a lot, the F1-score sits nearer the lower of the two — the harmonic mean punishes imbalance. In the example, the simple average is 73.3 but the F1-score is 72.7.
2. Red teaming for GenAI systems (AI-4.2.2)
The syllabus defines red teaming (RT) as "a systematic, often black-box form of fault attack that probes an AI-based system to identify harmful capabilities". It can cover the whole end-to-end system or only the model, and it is "especially critical for GenAI" because of the size of the attack space. In Sample Exam A v2.2 the red-teaming K3 item gives a scenario; live exams may phrase it differently, so learn the approach as a sequence you can apply to any scenario.
| Step | The approach typically involves (syllabus) |
|---|---|
| 1 | Assembling a diverse team of testers to cover a wide range of perspectives and attack vectors |
| 2 | Providing access to the AI-based system in a safe test environment |
| 3 | Prompting the system to identify vulnerabilities, through open-ended exploration or using checklists |
| 4 | Analyzing the identified failures to understand threats |
| 5 | Creating datasets from these threats to support mitigations and system improvements |
CT-AI Syllabus v2.0, section 4.2.2.
| Point the syllabus makes | Detail |
|---|---|
| When | Most effective before deployment, but after initial internal quality assessments are complete |
| Core activity | Interactive prompting, often multi-turn dialogues (e.g., 15-20 turns) |
| Scope | Security and safety, and also reliability, privacy, fairness, bias and misinformation |
| Security examples | Indirect prompt injection; malicious content hidden in documents used by RAG |
| Safety angle | Harmful outputs under ordinary use (e.g., unsafe medical advice) with no adversarial intent |
| Scaling it | Crowd-sourced prompts; an LLM generating attack prompts checked by another LLM; hybrid human + automated |
| Versus Blue Teaming | RT is proactive and pre-deployment focused; Blue Teaming is ongoing real-time monitoring and filtering in operation; RT insights improve Blue Teaming |
| Versus benchmarking | RT is adaptive (prompts change in response to outputs) and complements static benchmarking |
| Regulation | Increasingly a regulatory expectation, e.g., the EU AI Act |
CT-AI Syllabus v2.0, section 4.2.2.
3. Dataset constraint testing (AI-5.1.5)
Dataset constraint testing "verifies whether the data in a dataset adheres to predefined rules or constraints", much as a database schema constrains a table. The syllabus describes two scopes: single-value constraints on one attribute in a single instance, and multi-value constraints on an attribute across several instances. It then calls a comparison constraint "a special form of multi-value constraint" that compares different values. Know the ten names and where each sits; the Sample Exam A v2.2 item for this objective describes a situation and asks you to apply the constraints, and a live exam may frame it differently.
| Family | Constraint | Tests that… |
|---|---|---|
| Single-value (one attribute, one instance) | Missing | no values or attributes are missing |
| Single-value (one attribute, one instance) | Range | a value is in a given range |
| Single-value (one attribute, one instance) | Type | a value matches the specified type (an integer attribute holds no string or real number) |
| Multi-value (one attribute across instances) | Sum | the sum equals, exceeds or does not exceed a value (syllabus: Formula 1 race points cannot exceed 102 and must exceed 50.5) |
| Multi-value (one attribute across instances) | Count | the count of non-null values equals, exceeds or does not exceed a value |
| Multi-value (one attribute across instances) | Duplicate | identical or near-identical values or instances stay within a limit (often zero) |
| Multi-value (one attribute across instances) | Useful | an attribute has some repeated values — an all-unique attribute (ID, timestamp) gives the model no pattern |
| Multi-value (one attribute across instances) | Outlier | statistical outliers are identified |
| Multi-value, special form: comparison (two attributes) | Greater Than | one attribute exceeds another (lines of code exceed lines of code with defects) |
| Multi-value, special form: comparison (two attributes) | Correlate | two attributes correlate (students at least 1.33 standard deviations above the mean have grade A) |
CT-AI Syllabus v2.0, section 5.1.5. The syllabus adds that at realistic dataset sizes this testing is automated in the data pipeline, reporting to data scientists (training data anomalies) or operations staff (operational data problems).
Easy confusions
- Useful vs Duplicate: Duplicate limits repeats; Useful flags an attribute with no repeats at all.
- Range vs Outlier: Range checks a fixed allowed interval on each value; Outlier looks at a value against the statistics of all the others.
- Count vs Missing: Missing is a single-value check on one instance; Count is a multi-value check on how many non-null values there are overall.
- Constraint testing vs label correctness: checking that labels are right (expert review, multiple annotation, model-loss analysis) is a different objective, AI-5.1.6.
Our study notes on the syllabus definitions.
4. Metamorphic testing (AI-6.1.5)
Metamorphic testing (MT) derives new follow-up test cases from a previously passed source test case, using a metamorphic relation (MR): a property of the required function that says how a change to the inputs should show up in the expected results. It is one of the syllabus's techniques for the test oracle problem, alongside others such as A/B testing, back-to-back testing and expert review — you check relationships between outputs, not their absolute values. The objective asks you to derive test cases "for a given scenario", so expect to be asked for a valid MR or follow-up test.
| MR property (syllabus) | Meaning | Illustration (ours unless marked) |
|---|---|---|
| Consistency | Outputs align across related inputs | Rephrasing a search query without changing its meaning returns the same top results |
| Monotonicity | Outputs change in a direction as the input changes | Syllabus: in a risk model predicting age at death, more cigarettes smoked should lower the prediction |
| Invariance | Outputs remain stable under perturbations | Slightly brightening a photo does not change the object an image classifier recognises |
Properties and the smoking example from CT-AI Syllabus v2.0, section 6.1.5; the other two illustrations are ours.
| MT is typically selected when (syllabus) |
|---|
| No reliable expected outputs exist, due to model opacity or data scale |
| The system is a black box |
| Relational properties, not absolute values, are enough for confidence |
What the syllabus says MT cannot do
- MT detects relational flaws but not all absolute errors, so combine it with other techniques.
- Incorrect or incomplete MRs lead to false confidence (for example, overlooking interactions between variables).
- MRs come from domain knowledge, requirements or domain properties (e.g., laws of physics), and can be validated by expert review, reference models and checking edge-case coverage.
- MT normally starts from a source test that passed; the hands-on exercise 6.1.6 reminds students of the limitation when passed source tests are unavailable.
CT-AI Syllabus v2.0, sections 6.1.5 and 6.1.6.
The numbers you will not have to compute
Section 6.1.3 on statistical testing contains a worked example full of figures. The syllabus says in a note that candidates "will not be required to derive or compute these values using statistical formulae in the exam", so understand what they illustrate (why sample size matters, and why sequential testing can stop early) rather than memorising them.
| Syllabus example (6.1.3) | Figure |
|---|---|
| Target 98% accuracy, margin of error ±4%, 95% confidence | 601 test cases, of which at least 589 must pass |
| Same target with sequential testing, if observed accuracy stays consistently high (e.g., no less than 98%) | the ±4% margin at 95% confidence is reached after about 170 test cases |
| Safety criterion: 99% reliability with 95% confidence | 299 test cases, all of which must pass |
CT-AI Syllabus v2.0, section 6.1.3. Examined at K2 (AI-6.1.3), not K3.
How much the K3 questions are worth to you
Our arithmetic on ISTQB's pass mark of 29 of 44. Each K3 question lost has to be made up with two one-point questions.
| K3 questions right | Points from K3 | One-point questions still needed |
|---|---|---|
| 4 | 8 | 21 of 36 (58%) |
| 3 | 6 | 23 of 36 (64%) |
| 2 | 4 | 25 of 36 (69%) |
| 1 | 2 | 27 of 36 (75%) |
| 0 | 0 | 29 of 36 (81%) |
Our arithmetic. Percentages rounded to whole numbers.
Practise it, don't just read it
The Chapter 3, 5 and 6 quizzes include questions on metrics, input data testing and metamorphic testing; use them after you have worked the calculator.
The ProSyllabus CT-AI library has 10 quizzes of 10 questions written on the v2.0 syllabus and covering all seven chapters: Introduction to AI (two quizzes), quality characteristics, machine learning (two: forms and workflow; metrics, neural networks and coverage), testing AI-based systems (two: test oracles and test levels; generative AI and red teaming), input data testing, and model testing plus ML development testing (two). They are our own questions, not ISTQB exam items, and they are single-answer only, while ISTQB's official Sample Exam A v2.2 also includes two select-two questions, so practise that format there.
How many K3 questions are in the CT-AI v2.0 exam?
Four, one each for AI-3.3.1, AI-4.2.2, AI-5.1.5 and AI-6.1.5. Each is worth two points, 8 of the 44 in total.
Is the F1-score between 0 and 1 or 0 and 100 in CT-AI?
The v2.0 syllabus multiplies precision and recall by 100% and says F1 has values ranging from 0 to 100. Use that scale in the exam.
What is the formula for precision in CT-AI?
Precision = TP / (TP + FP) × 100%. Recall = TP / (TP + FN) × 100%.
Why does the confusion-matrix layout matter?
The syllabus puts Predicted on the rows and Actual on the columns but warns the axes may be swapped. Reading it the wrong way round swaps FP and FN, which swaps precision and recall.
Is the F1-score the average of precision and recall?
No. It is the harmonic mean, 2PR / (P + R). With precision 80 and recall 66.666… (40 of 60), the plain average is 73.3 but the F1-score is 72.7.
What are the steps of red teaming in the CT-AI syllabus?
Assemble a diverse team, give it access in a safe test environment, prompt the system to find vulnerabilities, analyse the failures, and create datasets from the threats to support mitigations.
What is the difference between red teaming and blue teaming?
Red teaming is proactive and pre-deployment focused (the syllabus says it is most effective before deployment, after initial internal quality assessments); blue teaming is ongoing real-time monitoring and filtering of inputs to an operational system. Red-team findings improve blue-team filters.
What kinds of dataset constraint does the syllabus name?
Single-value (Missing, Range, Type) and multi-value (Sum, Count, Duplicate, Useful, Outlier). Comparison constraints (Greater Than, Correlate) are described as a special form of multi-value constraint.
What is a metamorphic relation?
A property of the required function that describes how a change to a test's inputs should change its expected results. Follow-up tests are derived from a passed source test using it.
Do I need to calculate statistical sample sizes in the exam?
No. The syllabus says candidates will not be required to derive or compute the section 6.1.3 values using statistical formulae.
Sources
- official — ISTQB — CT-AI Syllabus v2.0 (GA 17 Apr 2026, PDF)
- official — ISTQB — Exam Structures & Rules Tables v1.19 (19 Aug 2026, PDF)
- official — ISTQB — Exam Structures and Rules v1.2 (2 May 2025, PDF)
- official — ISTQB — CT-AI Sample Exam set A, Questions v2.2 (PDF)
- official — ISTQB — CT-AI Sample Exam set A, Answers v2.2 (PDF, released 20 Jul 2026)
- official — ISTQB — Certified Tester AI Testing (CT-AI) certification page (v2.0)
official = a document published by the conducting body. reported = a news or coaching site we could not check against an original. Where sources disagree this page says so rather than picking one.










