Applied Judgment Assessment
Applied skill assessment · analysts and engineers who compute figures in Python · browse the full catalogue

Data Analysis in Python Assessment with Executed Scripts for Analysts and EngineersName the base. Compute it. Then decide whether the number that came out can be sent.

Six analytical questions, each answered in three stages: name the grain, the base or the statistic the question needs; compute it in Python that is actually run, once on the file shown and once on a file you have not seen; then look at the figure that came out and say whether it can be sent, and why not. Twenty-three shorter items between them on the denominator, the summary statistic under a long tail, the weighting trap, the period boundary and the blank inside an average. Scored as a path with recovery credit, so naming the wrong base and then catching it is a different finding from naming the right base and mis-computing it, and the report says which happened.

55 minutes41 scored exercisesEvidence-keyed scoringGlobal · INR & USD

A work sample about the analysis, not about the code

The Data Analysis in Python Assessment with Executed Scripts is a fifty-five-minute work sample for analysts and engineers who turn questions into figures. Six questions are each answered in three stages, name, compute and check, with the Python run on a shown and a hidden file; the report draws your figure against the one the data supports.

Most coding assessments measure whether a program passes its tests. That is a fact about the code. The figures that go wrong in real reporting usually go wrong before any code is written: a rate on the wrong base, a mean where the question needed a median, a month compared as if it were whole, a blank counted as zero, or an average of group averages that reverses when the groups differ in size. This instrument measures those decisions, and uses executed Python only as the way the figure gets computed.

Every one of the six programs is run twice. The file shown in the exercise is built so that the right figure and the figure of the most often named wrong base coincide on it; the second file, which you do not see, is where they differ, because it carries the repeat customer, the ninety-hour ticket, the shifted mix of sites, the four-day week, the blank and the customer not yet thirty days old. A program that holds on the shown file and not the hidden one is, on this instrument, the signature of computing a base other than the one the question needed. Every expected output was produced by executing a reference solution on both files before the exercise shipped.

The stages are scored as a path, and that is the point of the design. At the compute stage the best reachable value depends on what you named at the first stage: from a wrong base whose figure coincides on the shown file, the best you can do is pass the shown run, and doing exactly that earns full credit at the node. So a reader who names the wrong base, computes that base correctly, and then catches it at the check stage scores poorly at the first stage and well at the other two, and the report says exactly that in words. Without that credit one wrong name would be scored three times, which is a fact about the scoring and not about the analyst.

It is not the data-cleaning work sample in this catalogue, whose task is cleaning a file to a stated contract, and it is not the SQL work sample, whose task is writing the query to the right grain. Here the decision is about the analysis itself: whether the figure is the right figure and whether it can be sent. The construct is anchored on the data analyst occupation in ESCO, the European Commission's published classification of skills and occupations, and its skill statements on performing data analysis and interpreting data; ESCO is a free, published classification and no endorsement is implied.

The report is a dumbbell of difference. One row per question: the figure your program produced and the figure the data supports, two marks joined by a line, on that question's own scale, sorted by the size of the gap so the row to look at is the first one. A program that did not run, or was not attempted, is drawn with its gap absent and named at full size, never as a gap of zero. Beside each row the three stages are drawn as three marked cells with a word in each, and the declared table of what was reachable from each position is printed under it. A second dumbbell draws four competencies against the ordinary respondent with the 68 and 95 per cent bands on the chart.

Six executed programs, twelve single-choice items, eight select-every-line items, six true or false claims, five match-the-following items and four ordering items: forty-one in all, Python 3 with the standard library only, pandas not available. Every printed figure carries its band, every refusal is printed where the figure would have gone, and no percentile appears anywhere.

Six figures drawn against the data, four competencies against the ordinary respondent, and the stage at which each figure went right or wrong:
Grain and denominatorThe right summaryAbsent and partialChecking before reporting

What you walk away with

The figure your program produced against the figure the data supports

One row per question, two marks joined by a line on that question's own scale, sorted by the gap so the first row is the one to look at. A program that held on both files is drawn as two marks coinciding; a program that matched only the shown file is drawn against the figure the data supports, with the report saying the produced figure is inferred from the run pattern.

Which stage carried and which did not

Three marked cells beside every row: named the base or a different one, held on both files or the shown one only, caught it at the check or missed it. Naming the wrong base and then catching it reads as recovered; naming the right base and mis-computing it reads as a slip at compute. Those are different findings with different advice, and the report keeps them apart.

Four competencies with their bands on the chart

Grain and denominator, the right summary, absent and partial, and checking before reporting, each drawn against the ordinary respondent with the 68 per cent band as a solid inner block and the 95 per cent band as an outlined box, both written in words. A competency too thin to carry a figure carries a placement word and its band, and the reliability table says why.

A program that could not be run is never a wrong answer

Where the execution service cannot be reached, the row is drawn with its gap absent and DID NOT RUN printed at full size in its place. The node leaves both sides of its ratio and its chance term, and nothing on the page counts it against you.

The strength and what it costs, and one change

The competency furthest above the ordinary respondent, paired with the cost of having it in abundance: an analyst who always asks what the base is can stall on a definition nobody needs. Then one boxed if-then sentence naming the moment a figure is being decided and the small habit that moves it, with the movement a re-sitting in three months would have to clear.

Zero, declared

Zero on every figure is the ordinary respondent: somebody choosing at the declared share of respondents on each option, line, match and position, and passing the executed runs at the declared rate. Those shares are authored assumptions stated on the page, never uniform, and observed data replaces them.

Inside your report

Illustrative sample — your report is generated from your own responses and your own executed programs.

The figure your program produced against the figure the data supports, largest gap first
● filled circle = the data supports · □ open square = your program · a row with no comparable figure carries its reason at full size
1. Typical hours to resolve, by queue the summary statistic
Recovered
● data supports 3.0 hours□ your program 20.4 hours, 580 per cent apart024.1 hours (this case's own scale)
2. Complaint rate by branch the denominator
Slipped at compute
● data supports 0.333□ your program 0.500, 50 per cent apart00.590 (this case's own scale)
3. Orders per day, by week the period boundary
Incomplete
4. Pass rate by period the weighting
Unchecked
● data supports 0.827 — □ your program: the same, matched00.976 (this case's own scale)
Tabular fallback for the rows above.
CaseData supportsYour programGap
Typical hours to resolve, by queue3.0 hours20.4 hours580 per cent apart
Complaint rate by branch0.3330.50050 per cent apart
Orders per day, by week44.5 per day—did not run
Pass rate by period0.8270.827matched

How to read it: a program that held on the shown file and not the hidden one produced, by construction, the figure of the base whose output coincides on the shown file; that figure is drawn as the open square and the report says it is inferred from the run pattern. The first row is the one to look at, because the rows are sorted by the gap.

The three stages beside a row, and the four competencies against the ordinary respondent
Typical hours to resolve, by queue Recovered
1. Name it
named a different base
node 0.33
2. Compute it
held on the shown file only
node 1.00
3. Check it
caught it
node 1.00

You named the wrong base, computed that base correctly, and then caught it at the check. That scores poorly at the first stage and well at the other two, and the report says so: the analysis was wrong once and you would not have sent it.

Declared table for this case: best reachable value at each later stage, given what was named. ▶ marks the pick.
Named at stage 1Best at computeBest at check
▶ The mean of the hours, since it uses every ticket0.51.0
The middle value once the tickets are sorted (the keyed base)1.01.0
The value that occurs most often1.01.0
The mean after dropping the slowest tenth0.51.0
● you · □ the ordinary respondent at 0 · solid inner block = 68% band · outlined box = 95% band · ▲ above · ▬ near · ▼ below · ■ withheld
1. Grain and denominator
Below the ordinary respondent
□ ordinary 0● you -31−100+100
2. The right summary
Above the ordinary respondent
□ ordinary 0● you +24−100+100
3. Absent and partial
Near the ordinary respondent
□ ordinary 0● you +9−100+100
4. Checking before reporting
Withheld
Withheld, in the place the scale would have gone

Only 3 of 10 nodes in this competency were answered, under the 4 a placement needs. Nothing is printed for it.

Tabular fallback for the four scales.
CompetencyYou68%95%Placement
Grain and denominator-31-41 to -21-51 to -11Below the ordinary respondent
The right summary+2415 to 336 to 42Above the ordinary respondent
Absent and partial+90 to 18-9 to 27Near the ordinary respondent
Checking before reporting———Withheld

How to read it: zero is the ordinary respondent, computed from the declared share of respondents on every option, line, match and position and the declared pass rate on the executed programs. Never a percentile. A competency too thin to carry a figure is withheld in the place its scale would have gone.

Built for

  • Analysts and analytics engineers who compute rates, averages and shares in Python and want to know which of their figures would not have held on the next file
  • Hiring managers screening for the analysis behind the code: the base, the statistic, the treatment of blanks and partial periods, and the check before a figure is sent
  • Data teams that want the same six cases sat by everyone, with the stage at which each figure went wrong printed rather than a pass count
  • Anybody who has sent a number that turned out to rest on two orders, one ninety-hour ticket or a four-day week, and would rather find that out here

Find out which of your figures would have held on the file you did not see

41 items across six formats · 6 executed Python programs · about 55 minutes · ₹999 in India inclusive of GST, or US$9.99 elsewhere · 33 credits for organisations. One report: six figures against the data, the stage each one went right or wrong, four competencies with their bands, and one change to make on the next figure you compute.

₹999 (incl. GST) · assessment and full report, nothing further to pay

Buy this assessment

No account needed to buy. Your name and email identify the purchase and your receipt is sent to that address.

Secure Razorpay payment · ₹999 includes 18% GST

Bought this already and lost the tab? Sign in and enter your purchase code under Claim a purchase on your dashboard.

Secure checkout · INR & USDFull report immediately after submission

Frequently asked questions

Is this a Python coding test?

No, and the report says so. Six short programs are run in Python 3 with the standard library only, and a program that does not hold is reported as a figure that did not hold, not as a fault in your code. What is measured is the analysis: the grain and the denominator the question needs, the statistic that answers it on this distribution, the treatment of missing values and partial periods inside the figure, and the decision whether the figure that came out can be sent. Library breadth, software engineering, typing speed and the number of Python functions you can name are not measured.

How is this different from the data-cleaning and SQL work samples in the same catalogue?

All three execute your work, and that is where the likeness ends. The data-cleaning work sample asks you to clean a file to a stated contract; the SQL work sample asks you to write a query to the right grain. This one asks you to choose the base, the statistic and the treatment of blanks for a figure that the exercise deliberately does not specify, compute it, and then say whether the number that came out can be sent. The decision is about the analysis, and the code is only how it is computed.

What does recovery credit mean, and why does the scoring need it?

Each question is three stages, name, compute and check, scored as nodes on a path. At the compute stage the best value you can reach depends on what you named: from a wrong base whose figure coincides with the right one on the shown file, the best you can do is pass the shown run and fail the hidden one, and doing exactly that scores full credit at that node. So naming the wrong base, computing it correctly and catching it at the check scores poorly at the first stage and well at the other two, and the report says so in words. Without that credit one wrong name would be scored three times, which would be a fact about the scoring rather than about you. The declared table for each question is printed on the report.

What happens if my program cannot be run?

The row is drawn with its gap absent and DID NOT RUN printed at full size where the figure would have gone. The node leaves both sides of its ratio and its chance term, and the headline is computed over the stages that were answered and run. A program that ran and matched neither file is a different thing and is printed as such. Zero test cases passed with no execution is nobody running it, never a wrong answer.

How long does it take, what does it cost, and what is on the report?

Forty-one items across six formats, with a fifty-five-minute time limit: six executed Python programs, twelve single-choice items, eight select-every-line items, six true or false claims, five match-the-following items and four ordering items. The report costs ₹999 in India inclusive of GST, or US$9.99 elsewhere, and 33 credits for organisations. It draws each of the six figures your programs produced against the figure the data supports, sorted by the size of the gap, with the three stages of each question marked beside it; four competencies against the ordinary respondent with their 68 and 95 per cent bands; the strength paired with its cost; one if-then change with the movement a re-sitting in three months would have to clear; and every refusal printed where the figure would have gone. No percentile appears anywhere.

One of the AssessAll applied-judgment assessments

Each one takes a single capability, puts you inside the situations where it is actually tested, and scores your choices against published evidence — with a report designed for that capability alone, not a template. They span hiring, compliance, education, operations and personal skill.

Browse the catalogue →

Methodology: Forty-one original items across six formats: six executed Python programs, twelve single-choice items, eight select-every-line items, six true or false claims, five match-the-following items and four ordering items. CONSTRUCT STATEMENT: This measures whether an analyst can turn an analytical question into a computed figure that holds: choosing the grain and the denominator the question needs, the statistic that answers it on this distribution, the treatment of missing values and partial periods inside it, computing it in Python that actually runs, and deciding whether the figure that came out can be reported. It does not measure library breadth, software engineering, typing speed, or how many Python functions the reader can name. It is neither the data-cleaning work sample in this catalogue, whose task is cleaning a file to a stated contract, nor the SQL work sample, whose task is writing a query to the right grain: here the decision is about the analysis itself, whether the figure is the right figure and whether it can be sent, and the code is only how it is computed. DECLARED RESPONSE INSTRUCTION: knowledge and demonstrated skill throughout; every item is keyed and the six programs are executed; nothing is self-report. FRAMEWORK ANCHOR: ESCO, the European Skills, Competences, Qualifications and Occupations classification published by the European Commission, in particular the data analyst occupation and its skill statements on performing data analysis, analysing data and interpreting current data; cited as a published classification available under a free licence, with no endorsement implied. EXECUTION: each of the six programs is run twice, once on the input printed in the exercise and once on a second input the reader has not seen. The shown input is constructed so that the keyed figure and the figure of the most often named wrong base coincide on it, and the hidden input is constructed so that they differ; the hidden input carries the repeat customer, the ninety-hour ticket, the shifted mix of sites, the four-day week with a zero-order day, the blank and the all-blank region, and the customer not yet thirty days old. Every expected output was produced by running a reference solution in Python 3 with the standard library on both inputs. Where the execution service cannot be reached, the node is reported as not run and leaves both sides of its ratio rather than scoring zero, and the report says so on the page. SCORING DESIGN: C8, a path-value trajectory with recovery credit. Each case is three nodes, name it, compute it, check it, and every node scores (v_chosen - v_worst_available) / (v_best_available - v_worst_available). At the compute stage the best reachable value depends on the base named at stage 1: a program written to a base whose figure coincides with the keyed figure on the shown input can pass the shown run and not the hidden one, so from that position the best reachable value is 0.5 and computing the named base correctly scores 1.0 at that node. That normalisation is the recovery credit; without it one wrong name would be scored three times. Later stages are weighted more, 1 : 1.5 : 2, because a wrong name that is caught costs nothing downstream and a wrong figure that is sent costs everything. The total is chance-corrected against a respondent choosing at the declared marginals at every choice node and passing the executed runs at the declared execution prior; zero is the ordinary respondent. An unanswered node leaves numerator, denominator and chance term; an empty sitting scores exactly zero and is not reportable. Reliability is McDonald's omega in its Spearman-Brown form from the answered node count and an assumed average inter-item correlation of .30, printed as an assumption; a competency with fewer than eight answered nodes or omega under .70 unrounded carries a placement word and no figure. Every printed figure carries its SEM and its 68 and 95 per cent bands from a declared, modelled standard deviation. REPORT DESIGN: D4, a dumbbell of difference with a recovery column. One row per case, the figure the reader's program produced against the figure the data supports, dots encoded by shape and colour and labelled in words, sorted by relative gap largest first; a row whose program did not run or was not attempted is drawn with the gap absent and named, never as a gap of zero. Beside each row the three stages as three marked cells with a word in each, so that naming the wrong base and then catching it is a different finding from naming the right base and mis-computing it. A second dumbbell block draws the four competencies against the ordinary respondent with both bands. SOURCES, drawn across the literature rather than from one framework: ESCO (European Commission), the data analyst occupation and its data-analysis skill statements; Simpson, The interpretation of interaction in contingency tables (Journal of the Royal Statistical Society B, 1951), and Bickel, Hammel and O'Connell, Sex bias in graduate admissions: data from Berkeley (Science, 1975), for the reversal an unweighted mix of groups produces; Wainer, The most dangerous equation (American Scientist, 2007), for extremes appearing in the smallest groups; Tukey, Exploratory Data Analysis (1977), for the median and resistant summaries on long-tailed data; Hyndman and Fan, Sample quantiles in statistical packages (The American Statistician, 1996), for the definition of the median on an even count; Huff, How to Lie with Statistics (1954), for the average that is not the average the reader assumes; Rubin, Inference and missing data (Biometrika, 1976), and Little and Rubin, Statistical Analysis with Missing Data (2002), for why a blank is not a zero and why the mechanism of missingness matters; Hyndman and Athanasopoulos, Forecasting: Principles and Practice (2021), on calendar adjustment and partial periods; Wickham, Tidy data (Journal of Statistical Software, 2014), for one row per observation and the grain of a table; Broman and Woo, Data organization in spreadsheets (The American Statistician, 2018), for blanks, zeros and the difference between them; Haladyna, Downing and Rodriguez, A review of multiple-choice item-writing guidelines (Applied Measurement in Education, 2002), for cue control; McDonald, Test Theory: A Unified Treatment (1999), for omega; Jacobson and Truax, Clinical significance (Journal of Consulting and Clinical Psychology, 1991), for the Reliable Change Index; and Gollwitzer and Sheeran, Implementation intentions and goal achievement (Advances in Experimental Social Psychology, 2006), for the if-then action. The construct is public and taught; every item, every dataset and every expected output is an original work written and executed for this instrument. No commercial coding-assessment product, live-coding environment or simulation product is named or used, no competitor is named anywhere, and the instrument is not affiliated with or endorsed by the European Commission or any author named above.