All blueprints
Assessment blueprint · United Arab Emirates, Saudi Arabia & Qatar

What Should a Capability Baseline for an Emiratisation or Saudisation Development Cohort Actually Contain?

A capability baseline is the written specification of what a development cohort sits before a programme begins: which modules run, how many scored items each contributes, which competencies each reports, and what the result is allowed to decide. This one runs five live modules — 113 scored items, 112 minutes, 25 credits, US$12.50 a participant — and sets no pass mark.

Role: Emirati and Saudi national graduates and early-career joiners entering a private-sector nationalisation development programme

Module item counts and time limits read from the live module definitions. External facts on this page verified 16 September 2026.

The specification in short

  • ·Five live modules, 113 scored items, 112 minutes of test time, 25 credits — US$12.50 a participant at US$0.50 a credit, one sitting, whole cohort.
  • ·A 200-person intake is US$2,500. Run as a paired baseline and end-of-programme re-measure it is two sittings, US$25.00 a participant, US$5,000 for the cohort.
  • ·There is no stage 2 and no funnel saving. Every participant sits everything, so the cost is linear — which is what a development design costs and what a hiring design is engineered to avoid.
  • ·This blueprint sets no pass mark. All five modules ship an explicit passing_score of 0, and the platform treats an explicit 0 as no pass mark at all rather than as a bar everyone clears, so no credential is issued.
  • ·Counted across the shipped seed files on 16 September 2026: 256 of 304 assessments carry an explicit passing_score of 0, 40 carry 70 and 8 carry 60. Where the key is never set, the platform applies a legacy default of 70 — a number nobody chose for your cohort.
  • ·Nine distinct competencies are reported and all nine have a published definition on this site. That is not a boast about this page; it is a consequence of every module coming from one catalogue, and the previous blueprint in this gallery managed five out of ten.
  • ·Cognitive Speed is the most heavily measured construct here — 49 of the 113 items, across two modules, under two different display labels and one underlying slug. Nobody chose that, and a development plan cannot act on it.
  • ·The recommended weight is 1 on every module, which is the platform default. This is the only blueprint in this gallery where the default is the right configuration, and the reason is that the composite is not the output.
  • ·Four of the seven components named for NAFIS on the UAE Government's official portal are development measures rather than hiring measures — the Talent Program, Apprentice Program Support, On-the-job Training Support and Career Counselling — alongside the Emirati Salary Support Scheme, the Pension Program and the Child Allowance Scheme. Read 16 September 2026.
  • ·Checked at source on 16 September 2026: of three assessment and training-needs platforms selling into this work, one publishes a plan price and none publishes how many participants the price buys.
  • ·Every item count, time limit and credit price on this page is read from the live module definition in the repository, not estimated.

The blueprint

Weight is out of 5. It is the share of the combined score each module carries, and it is a decision — see the section below on what happens if you do not make it.

ModuleItemsMinutesCompetencies reportedWeightCredits
One sitting, sat by the whole cohort — there is no second stage, because nobody is being filtered out
Global English Proficiency

Corporate work in Dubai, Riyadh and Doha is conducted in English between colleagues who share no other language, and written English is where a graduate's first six months are most visibly graded. This runs first because it is the module whose result most often changes what the programme actually schedules in month one.

2825Verbal Reasoning · Attention to Detail15
Cognitive Agility & Learning Potential

How fast someone works out an unfamiliar rule, and what happens when the rule changes underneath them. In a development cohort this is the module that says how much scaffolding a person needs, not whether they belong here. Read its bands as a pace decision for the programme designer.

2422Abstract Reasoning · Logical Reasoning · Cognitive Speed15
Spreadsheet & Data Literacy

Absolute references, lookup semantics, format against value, and the single-column sort that quietly destroys a dataset. Spreadsheet competence is the most commonly assumed and least commonly taught skill in a graduate intake, and it is the one a programme can close in a fortnight once it knows who needs it.

2022Data Interpretation · Technology & Programming Logic15
Computer & Digital Fundamentals

Files and formats, operating systems, cloud concepts, networks, email mechanics. Every digital-transformation programme in the Gulf assumes this layer is already there; this is the module that finds out whether it is, per person, before anyone is put in front of a dashboard.

2518Technology & Programming Logic · Cognitive Speed & Agility15
Workplace Judgment — Situational Test

An error already sent to a client, two senior instructions that contradict each other, a Friday-evening discovery. These are the situations a first job produces in month two and that nobody is briefed on. The bands here are a mentoring agenda, and they are the part of the baseline a line manager will actually read.

1625Workplace Judgment · Critical Reasoning15
Whole blueprint113112US$12.50 per candidate at US$0.50 a credit525

Sample items are never published from a live module. Every item count above is the number of scored items in the module as configured, not an estimate.

What the blueprint covers

9 competencies across 5 modules, counted by the competency the module reports rather than by the label it prints. A competency carried by two modules is measured with more items and less error than one carried by a single module — which is a fact about your score, not about the candidate.

  • Cognitive Speed2 modules · 49 items
  • Technology & Programming Logic2 modules · 45 items
  • Attention to Detail1 module · 28 items
  • Verbal Reasoning1 module · 28 items
  • Abstract Reasoning1 module · 24 items
  • Logical Reasoning1 module · 24 items
  • Data Interpretation1 module · 20 items
  • Critical Reasoning1 module · 16 items
  • Workplace Judgment1 module · 16 items

What it does not cover

  • ·Arabic. Every module in this blueprint is in English. If the programme is delivered in Arabic, or if the sponsoring authority requires the assessment itself to be available in Arabic, this is the wrong specification and re-weighting does not fix it.
  • ·Spoken English. All five modules are written. A participant who writes well and freezes in a meeting will look fine here, and client-facing readiness is a different measurement on a different scale.
  • ·Technical or role-specific knowledge. Nothing here tests engineering, finance, petrochemical operations or any named system. A capability baseline says what a person can learn from and how fast; it does not say what they already know about your industry.
  • ·Motivation, intent to stay, and fit with the employer. Nothing in a 112-minute sitting predicts whether someone completes the programme. Completion in nationalisation cohorts is largely a function of line-manager quality, the visibility of a real job at the end, and how the first assignment is framed.
  • ·Leadership readiness. This is an early-career baseline. A leadership cohort needs a multi-rater design and a different instrument entirely.

The weight column, and what happens if you ignore it

The weight column is the centre of every other page in this gallery. On this one it is a decoy, and saying so is the most useful thing this page does.

AssessAll combines a multi-module sitting by averaging each module's percentage with a weight that defaults to 1.0. On the two hiring blueprints published here that default is the problem: it hands a 17-item module the same influence as a 40-item module on the number you rank people by. Here nobody is being ranked, so the default is harmless — and the composite it produces is the thing to ignore.

The reason is arithmetic rather than principle. A composite is an average, and an average is a statement about a person as a whole. A development plan is built from the gaps between modules, not from their mean. A participant who scores 82 on English and 41 on spreadsheet literacy has the same composite as one who scores 61 and 62, and the two of them need entirely different first months. Averaging is precisely the operation that destroys the information the programme was bought to find.

So the instruction is: configure the pack, leave the weights alone, and read the five module bands separately. If your programme dashboard shows a single number per participant, that number is the one field on the page with no use — and if anybody starts sorting the cohort by it, the baseline has quietly turned into a ranking.

The harder version of the same point, and it runs against us: if the composite is the field that gets read, you did not need a five-module blueprint. You needed one module and a conversation, and you should spend the other 20 credits on the re-measure instead.

How to configure it, honestly

  1. 1

    There is no one-click template and this page will not pretend there is. A blueprint is a specification you configure once; the configuration then repeats for every intake.

  2. 2

    Create one pack and add all five modules in the order in the table. The platform stores a pack as an ordered list, so the sequence you set is the sequence the participant sits. Do not split this into two packs — the two-stage design on the hiring blueprints exists to avoid spending on people you are about to reject, and you are not rejecting anyone.

  3. 3

    Leave the per-item weight unset. The default of 1.0 is correct here for the reason given above, and setting deliberate weights would imply the composite is meant to be read.

  4. 4

    Set no pass mark. Leave passing_score at the module default of 0 on every module. The platform reads an explicit 0 as no pass mark rather than as a bar everyone clears, suppresses all pass and fail framing in the result email, and issues no credential — which is the correct behaviour for an instrument whose output is a learning plan.

  5. 5

    Deliver by share link or QR to the cohort. Participants do not need accounts, which is what makes a single national sitting practical across intakes spread over Abu Dhabi, Dubai, Riyadh, Jeddah and Doha.

  6. 6

    Tell the cohort, in writing and before they sit, what the result will and will not be used for. A baseline sat under the impression that it is a selection test measures anxiety alongside everything else, and the people most likely to read it that way are exactly the early-career nationals the programme exists to develop.

  7. 7

    Book the re-measure at the same time you book the baseline. A paired baseline and outcome sitting is the only design here that supports a claim about the programme rather than about the people in it, and a re-measure that is scheduled after the results look disappointing is not evidence.

  8. 8

    Six months in, audit your own programme against this baseline: list everyone who was moved, deferred, streamed or dropped, and check whether any of those decisions traces back to a baseline score. If one does, the baseline became a selection instrument and nobody decided that it should. This is the step no vendor enjoys writing and it is the one that matters.

Four things this blueprint will not do for you

The construct this baseline measures most is the one a development plan cannot act on

Counted by competency slug rather than by display label, on 16 September 2026: Cognitive Speed is reported by two of the five modules across 49 of the 113 items, more than any other construct here. It appears once as Cognitive Speed and once as Cognitive Speed & Agility, which is two labels over one underlying competency, and a coverage table built from the display strings would report it as two things measured once each. Nobody chose this. It fell out of module selection, and it matters because you cannot write be faster at unfamiliar rules on an individual development plan. If the output of this baseline is a learning plan, either state plainly that two fifths of the evidence behind it is about pace, or substitute a module that does not report the speed band.

The pass mark you did not set is 70

Every module in this blueprint ships an explicit passing_score of 0, and the platform's submission handler treats that as no pass mark at all — not as a bar everyone clears — so all pass and fail framing is suppressed and no credential is issued. That is deliberate and it is documented in the handler. What the same line of code also does is apply a default of 70 to any assessment that never sets the key. Counted across the shipped seed files today, 256 of 304 set it explicitly to 0, 40 set 70 and 8 set 60, so the number is usually a decision — but if you add a module of your own to this pack and do not set it, your cohort inherits a legacy 70 that nobody chose. The handler also carries a fixed defect worth knowing about, because it is the kind that survives in other systems: an earlier version used a falsy check that silently coerced a deliberate 0 into 70.

A partial sitting still returns a number

The pack aggregate is computed over the segments a participant actually completed, not over the segments in the pack, so somebody who finished three modules of five returns a figure that looks exactly like a complete one. A completion flag is returned beside it. On a baseline this matters more than on a screen, because a development plan built from three modules will be silently missing two of its inputs and will read as though it were whole.

There is no AssessAll norm group behind any of these bands

A band here is a position against the instrument, not against Emirati graduates, Saudi graduates, or any other population. AssessAll publishes no norm group and no benchmark on these modules and does not claim one. That is a limit on what the baseline can say about a person and not a limit on what it can say about a programme: a paired baseline and re-measure compares the cohort with itself, which needs no external norm.

Comparing the national cohort with the expatriate cohort on one reported scale is the thing the standards forbid

Guideline SSI-2 of the International Test Commission's guidelines for translating and adapting tests states that scores should only be compared across populations when the level of invariance has been established on the scale on which scores are reported. A slide that puts national-cohort and expatriate-cohort averages side by side is doing exactly what that sentence rules out unless somebody has run the item-level check. AssessAll does not ship a differential item functioning analysis, and its adverse-impact monitor carries gender and age band only — no nationality dimension and no first-language dimension. The free calculator linked below will run the check on a cross-tab you export yourself, and it needs 200 people in the smaller group before it will classify anything.

The blueprint is a specification, not a validation study

Publishing item counts, durations and reported competencies tells you what the instrument is. It does not tell you that the instrument predicts success in your programme. Nobody has run that study on this cohort, here or anywhere else, and a vendor who tells you otherwise about a graduate development programme in the Gulf is describing a study that does not exist.

When not to use this blueprint

  • When the real decision is who gets a place. If the cohort is oversubscribed and the baseline will be used to choose, this is the wrong page and the wrong design — go and read the frontline hiring blueprint, set a cut score by a documented method, and call the thing what it is.
  • When the cohort is smaller than about ten people. Group patterns are the main reason to run a baseline rather than talk to everyone, and below ten there are no group patterns, only individuals you could have interviewed.
  • When the programme's curriculum is already fixed and cannot change in response to what the baseline finds. A diagnostic that nothing acts on is an expensive way to produce a file.
  • When the participant needs a certificate at the end. No pass mark means no credential is issued, by design. If the programme's output is a credential, that is a different instrument and a different configuration.
  • When the sponsoring authority requires assessment in Arabic.

Six questions to ask of any battery

These work on anyone's battery, ours included.

  1. 1

    How many scored items does each module contribute?

    A time limit is not a length. Two 20-minute modules can differ by 14 items, and the shorter one's items each move the composite further. Ask for the count, not the clock.

  2. 2

    What weight does each module carry in the combined score?

    If nobody can tell you, the answer is almost certainly equal weighting per module — which is a decision somebody made by not making it.

  3. 3

    Which competencies are covered twice, and which are covered once?

    A competency carried by two modules is measured with more items and less error than one carried by a single module. That is a fact about your score, not about the candidate.

  4. 4

    What does the composite do when a candidate does not finish?

    A partial sitting that still returns a number is the most common way two candidates get compared on different tests without anyone noticing.

  5. 5

    Where does the cut score come from?

    A bar is either set by a documented standard-setting method or it is set by someone's memory of a number. Ask which, and ask for the panel's spread.

  6. 6

    What is not on the blueprint at all?

    The competency nobody listed is the one the screen will not catch. Read the gaps before the coverage.

Questions buyers ask

What is a capability baseline for a nationalisation development cohort?

It is a single assessment sitting taken by everyone in the cohort before a development programme starts, used to decide what the programme teaches and in what order. It is not a selection test: nobody is excluded on the result. This blueprint specifies five live modules, 113 scored items, 112 minutes, 25 credits and US$12.50 a participant, with no pass mark.

Why does this blueprint set no pass mark?

Because nothing is being decided about the person. Every module here ships an explicit passing_score of 0, which the platform reads as no pass mark at all rather than as a bar everyone clears — pass and fail framing is suppressed in the result and no credential is issued. Counted across the shipped assessment files on 16 September 2026, 256 of 304 set it to 0 explicitly.

How much does a capability baseline cost per participant in US dollars?

On this blueprint, US$12.50 a participant for one sitting — 25 credits at US$0.50 a credit. A 200-person intake is US$2,500. A paired baseline and end-of-programme re-measure doubles it to US$25.00 a participant, or US$5,000 for the cohort, and that pairing is the only design here that supports a claim about the programme.

Should I rank the cohort by the combined score?

No. The combined score is an average across modules, and a development plan is built from the gaps between modules rather than from their mean. A participant scoring 82 on English and 41 on spreadsheet literacy has the same average as one scoring 61 and 62, and the two need different first months. Read the five module bands separately.

Can I compare the national cohort's scores with the expatriate cohort's?

Not safely, and not without work. The International Test Commission's guideline SSI-2 says scores should only be compared across populations once invariance has been established on the reported scale. AssessAll ships no differential item functioning analysis, and its adverse-impact monitor covers gender and age band only. Run the item-level check yourself before any slide puts the two averages side by side.

Do assessment vendors publish what a capability baseline costs?

Mostly not. Checked at source on 16 September 2026: TestGorilla publishes plan prices (Core at US$142 a month, US$1,704 billed annually) and a credit rate; Mercer Mettl's pricing page carries no number in any currency; and one specialist training-needs platform has a page titled Pricing that contains no price and routes to email. One of the three publishes a price, and none publishes how many participants the price actually buys.

Is a capability baseline part of Emiratisation or Saudisation compliance?

No. Nationalisation obligations are about employment, and no rule in the UAE, Saudi Arabia or Qatar requires an assessment. What the programmes do fund is development: four of the seven NAFIS components named on the UAE Government's official portal — the Talent Program, Apprentice Program Support, On-the-job Training Support and Career Counselling — are development measures rather than hiring measures. A baseline is useful because it tells that spending where to go, not because a regulator asks for it. Nothing on this page is legal advice and AssessAll holds no accreditation from any Gulf authority.

How does this differ from the hiring blueprints in this gallery?

Three ways. A hiring blueprint runs in two stages so the expensive modules never reach the whole applicant pool; this runs as one sitting because nobody is filtered out, so the cost is linear. A hiring blueprint needs deliberate weights because the composite is the ranking; here the composite is the field to ignore. And a hiring blueprint needs a cut score set by a documented method; this one has no pass mark at all.

SiddharthanFounder, AssessAll — Bodhih Training Solutions

Founder of AssessAll and of Bodhih Training Solutions, a corporate training company in Bangalore. Works on assessment design, scoring and reporting across hiring, L&D and certification programmes.

Last reviewed

This page describes assessment design, not law. Nothing here is legal advice, and AssessAll holds no accreditation or certification in Singapore or Malaysia.

Sources

Related

Other blueprints