What it costs you
Nothing. Cornell funds the build, the platform, and the analysis. There is no licence, no per-seat fee, and no commercial commitment attached to a partnership.
For Employers
We build and validate realistic work simulations that measure whether people can actually do their work well with AI in it — not whether they finished a course or say they feel confident. Partner organizations shape what we build, run it with their own teams, and see the results first.
No cost. No software to install. Your organization can stay anonymous.
Why this exists
Most organizations can report how many employees completed AI training and how many say they feel confident using AI tools. Neither number tells you whether anyone does better work. Course completion measures exposure and self-report measures comfort, and both leave leaders guessing about the thing they actually need to plan around.
The alternative most often proposed — a skills taxonomy, a certification, a battery of questions about AI — has the opposite problem. It is precise about categories and silent about behavior. Knowing that someone can define a hallucination does not tell you whether they check the output before it goes to a customer.
What is missing is evidence from the work itself. We put a person into a realistic task from their own function, with AI available, and observe what they do at each step: what they hand to the model and what they keep, how they specify the work, whether they verify against the source, and whether they take responsibility for what they deliver. They receive targeted feedback and attempt a second, comparable task. The change between the two attempts is a measurement in its own right — it separates a gap that closes with coaching from one that does not.
What we measure
Deciding when and how to use AI, describing the work clearly, evaluating what comes back, and taking responsibility for the result. Measured individually and across team workflows.
How quickly someone picks up a new approach, integrates feedback, and carries what they learned into the next task.
Working through the human parts AI does not remove: giving hard feedback, repairing after a mistake, and keeping a room safe enough that people say what they actually think.
Ways to partner
Not every organization is ready to run a pilot, and you do not need to be. Each of these is useful to us on its own, and each one is a reasonable place to stop.
1
About one hour, one conversation
We interview someone who knows a role well and turn what they describe into a work pattern — the sequence of decisions a real task actually involves. This is the raw material for everything else we build. You get early access to whatever we build from it, and no further obligation.
2
Two to three hours of subject-matter-expert time
We build a scenario around work your people actually do, with your documents, your constraints, and your definition of a good outcome. You review it, we revise it, and you keep access to the result. The competency framework and scoring stay constant; the content is yours.
3
About one hour per participant · eight to nine weeks end to end
Forty to sixty of your employees complete the simulation. Each one receives private feedback. You receive an aggregate report on where the capability sits, where the gaps are, and how much those gaps closed after a single round of feedback — plus a written interpretation from us and a live debrief.
4
Ongoing, for organizations willing to go further
The open research question is whether simulation performance predicts outcomes employers actually care about — ramp time, quality, retention, promotion readiness. Answering it requires partners willing to connect assessment results to outcomes over time. If that interests you, it is the most valuable thing you can do with us.
The exchange
Nothing. Cornell funds the build, the platform, and the analysis. There is no licence, no per-seat fee, and no commercial commitment attached to a partnership.
This is a research initiative, and the thing we cannot manufacture is work that is real. We ask for the subject-matter-expert time to make a scenario honest, permission to use de-identified data in academic research, and your candid reaction to what we produce — including when it turns out not to be useful. A null result is a finding we need.
Your people and your data
Participation is voluntary, carries no employment consequence, and both facts are disclosed to participants before they consent. All data collection runs under Cornell University IRB approval.
Individual results go to the individual. Your organization receives aggregated, de-identified results at the group or function level, with a minimum reporting group size of five and cross-tabulation limited so small groups cannot be reconstructed. Cornell holds identifiable data as the research data controller; it is not transferred to you.
We recommend against giving managers visibility into individual results, at least on a first run. The reason is practical rather than ideological: participants who believe their results reach their manager tend to perform to the instrument rather than work the way they normally would, which costs you both the accuracy of the measurement and the honesty of the feedback.
Every partnership is governed by a Data Use Agreement covering purpose, retention, security, and publication. Your organization is not named in any public material or publication unless you choose to be. Request our standard data terms →
Deliverables
A scored profile across each competency dimension, anchored to specific evidence quoted from their own work. Written feedback on what to do differently. Their own change between the first and second attempt.
Score distributions by competency and dimension. A dimension-by-function map of strengths and gaps. The growth signal — how much people improved after one round of feedback, and how many did. Behavioral indicators: who verified AI output against the source, who disclosed their AI use, who correctly chose not to use AI at all. Development recommendations tied to specific dimensions. A written interpretation from our team, and a live debrief.
On benchmarks. No credible external benchmark for these competencies exists yet, and we would rather say so than imply otherwise. What we offer is comparison within your own organization, comparison of each person against their own earlier attempt, and reference distributions from our university pilots. The instrument is built for feedback and development, not for a score that ranks you against an outside standard.
Fit
The simulations currently target desk-based work where AI tools sit inside the workflow — analysis, writing, planning, client-facing communication, and the judgment calls around them. We have built scenarios in and around finance, human resources, sales, operations, and technology functions.
We have not yet built for frontline and field roles where the work is physical and the AI touchpoint is narrower. It is a direction we find genuinely interesting and want partners for, but we would not want an organization to expect a deployable assessment there today.
Organizations of any size can partner with us. The pilot design needs roughly forty participants to produce a useful signal, but the first three ways in have no size requirement at all.
Timeline
| Weeks | What happens |
|---|---|
| 0 | Intro call. Agree scope, participant count, and how results will be shared |
| 1–3 | Data Use Agreement in review, in parallel with the co-design session and scenario build |
| 4 | Participant recruitment and internal communications |
| 5–6 | Participants complete the simulation. Each receives feedback immediately |
| 7–8 | Analysis; aggregate report and written interpretation delivered |
| 8–9 | Leadership debrief |
In our experience the legal review, not the design work, is what moves this date. Starting it in week 0 is the single best thing a partner can do to hold the schedule.
Questions
No, and we would push back on running it that way. It is a development instrument. Results are not tied to employment decisions, and we ask partners to say so explicitly when they invite people to participate.
No. Part of what the simulation measures is what someone does when a capable tool is available and they have to decide whether and how to use it. We do ask which tools your people actually have access to, so the scenario reflects their real environment rather than a generic one.
The environment is model-agnostic and we configure it to resemble the tooling your participants actually have.
Yes. Your organization is unnamed in all public materials and publications unless you opt in. Several of our partners prefer this and it changes nothing about what you receive.
For a pilot, roughly two hours from the executive sponsor, two to three hours from the subject-matter experts who help design the scenario, and three to four hours from whoever coordinates logistics. Participant time is about an hour each.
A full pilot needs around forty participants to produce a signal worth acting on. Below that, the first two ways in are a better fit and are equally useful to us.
That is genuinely open, and depends on what we find. Some partners run a second cohort, some extend to new functions, and some tell us it did not earn a place in their process. We want to know either way.
Tell us what your people do and where AI is changing it. We will tell you honestly whether we can help, and what it would take.
Rachel Slama, Associate Director — rslama@cornell.edu
Cornell University · Computing and Information Science Building 363 · Ithaca, NY