Key takeaways

  • ✓An AI literacy baseline is a structured assessment of what your people actually know and can do with AI today, before you spend a dollar on training.

  • ✓Skipping the baseline is the most common reason enterprise AI training fails to stick: teams sit through content pitched at the wrong level.

  • ✓A useful AI literacy assessment measures more than awareness. It covers confidence, current tool use, workflow fit, and attitude toward AI adoption.

  • ✓Assessment results should map directly to role clusters and training tiers, not a single all-staff program.

  • ✓The process does not need to be complex. A well-designed survey, a handful of team-level conversations, and a clear scoring framework are enough to get started.

What does 'AI literacy baseline' actually mean?

An AI literacy baseline is a snapshot of what your people currently know, believe, and can do with AI tools before any formal training begins. It tells you where the organisation actually stands, not where you hope it stands or where a vendor's pre-sales deck assumes it stands.

The word "literacy" is doing real work here. It does not mean everyone becomes a data scientist. It means people can read AI outputs critically, judge when to trust them, prompt tools effectively for their role, and understand enough about how models behave to avoid costly mistakes. A finance analyst approving an AI-generated summary needs different literacy than a procurement manager using Copilot to draft supplier briefs. The baseline captures both.

How is this different from AI readiness?

AI readiness is a broader organisational question: do you have the infrastructure, governance, and culture to adopt AI at scale? Literacy is the human layer inside that. You can have a perfectly configured Microsoft 365 tenant and a responsible AI policy written by lawyers, and still have 400 staff who do not know how to write a useful prompt or who believe every output is factually reliable.

The two concepts are related, but confusing them leads to a specific mistake: organisations invest in platforms and policies, then assume the people side will sort itself out. It rarely does.

Why "baseline" rather than just "survey"?

A survey collects opinions. A baseline measures capability. It gives you a reference point you can actually return to in six or twelve months and ask: did training move the needle? Without that starting point, any claim about training effectiveness is educated guesswork.

This distinction matters when you are presenting a business case internally or reporting to a board that wants to see return on a training budget. Numbers need a denominator, and the baseline is it.

Why skip the baseline and you waste the budget?

Training without a baseline is essentially guessing. You pick a program, schedule the cohorts, and hope the content lands. Sometimes it does. More often, you end up with a room where half the participants already know the material and the other half aren't ready for it yet.

That mismatch is expensive in two directions.

The first problem is over-training. Senior analysts or developers who already work with AI tools daily will sit through foundational content they covered months ago. They disengage, they stop attending, and the message that spreads back through the organisation is that the training wasn't worth their time. That reputational damage affects uptake for every program that follows.

The second problem is under-training, which is subtler and harder to recover from. A team that hasn't yet grasped what an AI tool actually does, or why they'd use it, will struggle with task-specific training before they're conceptually ready. They leave a workshop with techniques but no mental model for when to apply them. Adoption stays low, and the organisation concludes the tool is the problem when the real gap was sequencing.

The sequencing problem is often invisible

When training fails, the instinct is to blame the content or the vendor. The more common cause is that the training was pitched at the wrong level for the audience, because no one checked where the audience actually was.

There's a budget dimension too. Enterprise AI training isn't cheap, and neither is the time cost of pulling teams out of their work. If you're running the same program across a legal team, a marketing department, and a technology group without any sense of how differently those groups currently use AI, you're not being efficient. You're applying a uniform solution to a varied problem.

A proper AI literacy assessment takes that variation seriously. It surfaces the departments that need foundational awareness before anything else, the teams ready for role-specific capability building, and the individuals who could become internal champions if pointed in the right direction. That kind of segmentation is what turns a training budget into a training strategy.

The other risk of skipping the baseline is that you can't measure progress. If you don't know where people started, you have no honest way to show whether the training moved the needle. That matters when you're reporting back to a CFO or a board that wants to see return on a significant investment. Measuring ROI on enterprise AI training becomes a lot easier when you have a clear before-and-after picture, and that picture starts with the baseline.

How do you run an AI literacy assessment?

A practical AI literacy assessment has four moving parts: deciding who to include, choosing your methods, running the data collection, and setting a timeframe that doesn't exhaust your stakeholders. None of these need to be complicated.

Decide who is in scope

Resist the temptation to assess everyone at once. For a first baseline, sample across three groups: frontline staff who will use AI tools day-to-day, team leads and middle managers who will shape adoption, and a handful of senior leaders whose mental models of AI will influence budget and policy decisions. Aim for genuine representation across departments rather than a single business unit, or the picture you get will be too narrow to act on.

If your organisation has already identified specific functions as early AI adopters (finance, legal, customer service are common candidates), weight your sample toward those groups. You'll get more actionable data from the people who need to move first.

Choose your methods

No single method gives you the full picture. Use two or three in combination.

Self-assessment survey. The fastest and most scalable option. A well-designed survey of 15 to 20 questions can cover awareness, confidence, current usage habits, and perceived barriers in under ten minutes. The limitation is obvious: self-reported confidence is not the same as actual capability. People tend to overestimate what they know and underestimate what they don't. Use survey results as a directional signal, not a definitive score.

Structured interview or focus group. Talking to six to eight people across different roles for 30 minutes each surfaces things a survey never will: the workarounds people have invented, the fears they won't write down, the use cases they're already experimenting with informally. Focus groups work well for teams with similar roles; one-on-ones are better for senior stakeholders who are less likely to speak candidly in a group.

Task observation. If you want to understand capability rather than just confidence, ask a small group to complete a realistic task using an AI tool they have access to. Watch how they construct a prompt, how they evaluate the output, and what they do when the output is wrong. You'll learn more in 20 minutes of observation than from a dozen survey responses. This method requires more coordination, so reserve it for the roles where AI capability matters most.

A note on existing data

Before you design anything from scratch, check what you already have. Learning management system (LMS) completion rates for any AI-related content, help desk tickets about AI tools, and adoption metrics from tools like Microsoft Copilot can all give you a meaningful head start. Low Copilot adoption in a particular team, for example, is itself a data point about literacy gaps. There's also a broader question of AI readiness that goes beyond individual skill, which this AI readiness assessment framework covers in more depth.

How long does it take?

For a mid-sized organisation (200 to 2,000 employees), a well-scoped assessment typically runs across two to three weeks:

Phase

Activity

Timeframe

Design

Draft survey, recruit interviewees, align stakeholders

3-5 days

Data collection

Survey live, interviews and observations running

5-7 days

Analysis

Synthesise findings, identify patterns by role and department

3-5 days

Reporting

Write summary and initial recommendations

1-2 days

That's roughly three weeks end-to-end if you move deliberately. Organisations that try to compress this into a single week usually end up with too few interviewees and survey data they can't interpret confidently.

Keep the scope honest

An assessment that tries to measure everything across every team at once tends to produce data too aggregated to act on. A tighter scope, two or three methods, and a representative sample will give you clearer direction than a company-wide survey sent to 3,000 people with a 12% response rate.

One practical tip: pair the assessment with a genuine communication about why it's happening. Staff are more likely to respond honestly if they understand the results will shape training, not performance reviews. That framing also gets you better focus group participation and more candid interview responses.

What should an AI literacy assessment actually measure?

A useful AI literacy assessment covers four dimensions. Most organisations that have tried a quick survey find they only captured one or two of them, then wonder why the training didn't land.

Conceptual understanding is where most people start, and rightly so. Can someone explain what a large language model actually does? Do they understand why an AI tool might confidently produce an incorrect answer? Do they know the difference between a general-purpose assistant and a tool trained on proprietary data? You don't need engineers here. A finance manager doesn't need to know how a transformer architecture works, but they do need to understand that AI outputs require human judgement before they go into a board report.

Tool familiarity is more specific: which tools does this person already use, how often, and in what contexts? Someone who opens Microsoft Copilot every morning to draft meeting notes is in a very different place from a colleague who has logged in once and hasn't returned. This dimension often surfaces surprising gaps. Senior leaders sometimes have the lowest tool familiarity of anyone in the business, simply because no one thought to train them first.

Workflow application asks whether people can connect AI capabilities to their actual work. This is where literacy turns into fluency. Knowing that Copilot can summarise documents is conceptual. Knowing which of your team's three most time-consuming tasks it could meaningfully accelerate is application. Assessing this dimension usually requires either short scenario-based questions or a brief structured conversation, because it is contextual in a way that a multiple-choice survey cannot fully capture.

Attitude and confidence is the dimension that gets skipped most often, and it's the one that predicts adoption more reliably than the others. An employee who scores well on conceptual knowledge but feels anxious, sceptical, or quietly resistant will not use the tools consistently. Conversely, someone with genuine curiosity and low anxiety will experiment, build habits, and help colleagues along the way. Measuring this honestly means giving people a safe way to express doubt or concern, not just asking whether they feel "ready."

Four dimensions, not one

Conceptual understanding, tool familiarity, workflow application, and confidence each predict different outcomes. An assessment that only measures knowledge will miss the people who understand the tools but won't use them, and the people who are eager but need structured support.

A fifth area worth including for certain roles is data and ethics awareness: understanding what should and shouldn't go into a public AI tool, what the organisation's acceptable use policy actually says, and how to recognise outputs that warrant scrutiny. For teams handling sensitive data, this isn't optional.

If you want a more structured framework for thinking about organisational readiness across all of these dimensions, the AI readiness assessment guide covers the full picture, including how to map gaps at a team and business-unit level rather than just individually.

How do you turn assessment results into a training plan?

Assessment data is only useful if it changes what you do next. The goal is to move from a spreadsheet of scores to a set of cohorts, each with a clear training brief.

Start by grouping people, not by job title, but by what the data actually shows. A typical enterprise ends up with three or four bands:

  • Foundational (little or no working knowledge of AI tools, unclear on basic concepts like prompting or model limitations)

  • Functional (understands the basics, uses AI occasionally, but inconsistently and without much confidence)

  • Proficient (uses AI tools regularly in their work, can coach peers informally, ready for more advanced or role-specific content)

  • Advanced (strong technical or applied understanding, potential train-the-trainer candidates)

These bands matter because each one calls for a different response. Putting a proficient user through a foundations workshop wastes their time and yours. Sending a foundational learner straight to a prompt engineering deep-dive sets them up to disengage.

Matching cohorts to training formats

Once you have your bands, the format question becomes more tractable. A few principles that hold across most enterprise contexts:

Cohort

What they need

Format options

Foundational

Confidence, context, basic tool fluency

Facilitated workshop, short e-learning modules

Functional

Consistent habits, role-specific use cases

Role-specific workshop, guided practice sessions

Proficient

Depth, broader application, peer learning

Advanced workshop, community of practice

Advanced

Influence and scale

Train-the-trainer program, internal champions

For foundational and functional cohorts, in-person or live virtual delivery tends to work better than self-paced e-learning. People who are uncertain about AI benefit from being able to ask questions and see tools used in real time. The custom vs off-the-shelf decision is worth thinking through carefully here: off-the-shelf modules can cover general concepts, but they rarely address the specific tools, workflows and risk appetite your organisation actually has.

For proficient and advanced cohorts, the training need is less about hours in a classroom and more about depth and application. A well-run community of practice, combined with occasional expert-led sessions on emerging capability, often delivers more than another full-day workshop.

Use the data to prioritise, not to treat everyone at once

One of the more common mistakes at this stage is trying to run training for every cohort simultaneously. That stretches L&D resources, makes scheduling a nightmare, and usually means none of the cohorts gets the attention it needs.

A more workable approach is sequencing. Identify which cohort has the highest organisational impact if uplifted quickly. For most organisations, that is the functional group: they already have some exposure to AI tools, so the gap between where they are and where they could be is relatively small, and closing it tends to show results quickly. Visible early wins build internal momentum for the rest of the rollout.

Sequence by impact, not by size

Start with the cohort where a modest skills uplift produces the clearest business result. Quick, visible wins build the internal case for the broader program.

The L&D playbook for a global AI training rollout covers sequencing and prioritisation in more depth, including how to manage rollouts across multiple business units or geographies without losing coherence.

Finally, build a feedback loop into the plan from day one. Assessment scores taken before training and again at the 60 or 90-day mark give you evidence that the program is moving the dial, which matters both for internal reporting and for decisions about where to invest next.

Frequently asked questions

How long does an AI literacy assessment take to run?

A well-designed assessment takes most employees between 15 and 25 minutes to complete. The larger time investment is in the planning phase before you launch: mapping the roles you want to assess, agreeing on what good looks like for each, and building the communication that gets people to take it seriously rather than clicking through. Budget two to four weeks for that groundwork, and you will get cleaner data from the assessment itself.

Do we need specialist tools, or can we run this in-house?

You can run a credible baseline with tools your organisation already has. A structured survey in Microsoft Forms or Google Forms, combined with a short practical task graded against a rubric, is enough for most enterprises. Dedicated skills intelligence platforms exist if you want automated scoring at scale, but they are not a prerequisite. The quality of the questions matters far more than the sophistication of the platform.

How often should we reassess?

Reassess every six to twelve months. AI tools and workflows change quickly enough that a baseline from eighteen months ago is largely obsolete. Many organisations tie reassessments to major product updates, for example when a new Copilot capability rolls out, or to the completion of a training cohort so they can measure the shift.

What if employees feel threatened by being assessed?

Frame the assessment as a planning tool for the organisation, not a performance review for individuals. Share results at team or department level rather than surfacing individual scores to managers. When people understand that the output is a training plan designed to support them, resistance drops considerably. Positioning AI fluency as a team capability rather than an individual grade helps reinforce that framing.

Can we use the same assessment across very different roles?

Not reliably. A single generic assessment will produce data that is too blurry to act on. You do not need a completely separate instrument for every job family, but you do need role-specific questions or scenarios in the practical component. An accounts payable officer and a data analyst both need AI literacy; what that means in their day-to-day work is different enough that the same rubric will miss the gaps that matter for each.

Ready to map where your teams actually stand?

Most AI training programs that stall do so before the first session, not because of the training itself, but because no one ran the diagnostic first. Skipping the baseline is the single most common reason budget gets spent and behaviour doesn't change.

A short discovery conversation with Better People covers what a useful AI literacy assessment looks like for your organisation, how to structure it across roles and business units, and what a realistic training plan looks like once the results are in.

Not sure where your teams actually sit on AI fluency?

We'll walk through what an AI literacy assessment covers for your context, and what a structured training plan typically looks like from there.

Book a 30-minute discovery call →