
AI for Engineering Productivity
A practical playbook for running an AI pilot in a mechanical engineering team: cohort sizing, baseline metrics, workload choice, and a 90 day decision loop.
·
⏱
9 min read

Dr. Maor Farid
Maor Farid is the Co-Founder and CEO of Leo AI, the first AI platform purpose-built for mechanical engineers. He holds a PhD in Mechanical Engineering and completed postdoctoral research at MIT as a Fulbright fellow. A Forbes 30 Under 30 honoree and former AI researcher and Mechanical Engineer in an elite military intelligence, Maor leads Leo AI's mission to transform how engineering teams design better products faster.

BOTTOM LINE
A pilot exists to produce a decision, not an impression. Size it at five to 20 users so you get enough weekly signal to see a pattern while staying small enough to support and interview directly. Compose the cohort across seniority and function rather than by who volunteers. Baseline three to five metrics before anyone logs in, write the threshold next to each one, and fix the decision date in the calendar. Point the pilot at one frequent, measurable workload against real data rather than a sanitised sample, and resist adding a second workload once it is running. Then hold to 90 days and write the result down. Done that way, a pilot that fails is almost as useful as one that succeeds, because you know why.
Most engineering teams do not fail at AI because they chose the wrong product. They fail because nobody designed the pilot. Somebody gets a licence, a few engineers try it for two weeks, opinions form, and the whole thing quietly ends with no one able to say what actually changed.
The MIT NANDA study The GenAI Divide: State of AI in Business 2025 puts numbers on how common that ending is. Sixty percent of organizations evaluated enterprise grade AI systems, 20 percent reached a pilot, and 5 percent reached production. The distance between that second number and the third is where most engineering AI budgets disappear.
A pilot is not a trial period. It is an experiment with a cohort, a baseline, a threshold, and a date on which somebody decides. This playbook covers how to build one for a mechanical engineering team of five to 20 users, and how to make the outcome defensible either way.
Why engineering AI pilots stall before production
The failures are repetitive enough to list. Across the pilots that never convert, four patterns account for nearly all of them.
No baseline. Nobody recorded how long the task took beforehand, so every result is anecdote. An engineer saying it felt faster is not evidence a finance team can act on.
The tool never touched real data. Pilots run on a sanitised folder of sample models or on public reference material. That tells you the interface works. It tells you nothing about whether the system can find anything in a vault with 15 years of inconsistent naming in it.
No owner with authority. When a pilot belongs to everyone, the decision belongs to no one. Pilots without a named owner tend to expire rather than conclude.
The system does not learn. The NANDA research identifies this as the root cause rather than a symptom: most deployed systems do not retain feedback, adapt to local context, or improve with use, so week eight looks exactly like week one.
McKinsey has a name for the resulting condition, pilot purgatory, in which use cases keep being trialled but none receives the data access, process change, or integration it would need to reach production. The engineering version has a specific texture: the tool works in the demo, works on the sample set, and falls over the first time somebody asks it about the fixture from the 2019 programme that only one person remembers.
Every item on that list is a design problem, which means every item is fixable before the pilot starts.
IN PRACTICE
It integrates directly with PLM and existing workflows, making past designs, standards, and calculations instantly available. The result is fewer errors, faster decision-making, and a more consistent process across teams.
- Sergey G., Board Member
Size the pilot for signal: the 5 to 20 user range
Pilot size is usually decided by whoever has spare licences. It should be decided by how much signal you need.
Two or three users is too few. One strong personality dominates the read, the weekly query volume is too low to show a pattern, and you only see the handful of question types those individuals happen to ask. A single sceptic or a single enthusiast can swing the conclusion, and neither reflects how the wider team will behave.
Fifty users is too many for a first pass. Support load grows faster than insight, the integration work expands to satisfy edge cases you have not yet justified, and procurement gets involved before you hold any evidence. You end up defending a rollout instead of testing an idea.
Five to 20 users is the working range because it is the smallest group that produces enough weekly sessions to see repeated behaviour, while still being small enough that one person can support all of them directly and speak to each of them every week. That second property matters more than it sounds. The most useful pilot data is qualitative and arrives in conversation, and it stops arriving as soon as the cohort outgrows a single reviewer.
Composition matters as much as headcount. A cohort worth learning from usually contains all of the following:
A senior engineer who holds a large share of the team's undocumented history, because they are the hardest user to impress and the most valuable one to convince.
One or two engineers in their first year, who will surface every gap in how findable your existing information is.
A designer or drafter whose work is dominated by reuse and revision rather than new geometry.
Somebody from manufacturing engineering, so that manufacturability and cost questions enter the test set rather than only design questions.
The PDM or PLM administrator, who is the only person who can tell you quickly whether an integration issue is the tool or the vault.
Define success metrics before the first login
The single highest value hour in a pilot happens before anyone has access. In that hour you write down what you are measuring, what the current value is, and what number would justify moving to production.
Three to five metrics is the right count. Fewer and you cannot distinguish a real gain from a coincidence. More and nobody maintains them past week three. For a mechanical engineering team the candidates that reliably produce clean numbers are:
Time to locate an existing part, drawing, or calculation. Time a set of 10 realistic retrieval tasks before access, then repeat the identical set at week eight.
Part reuse rate. The share of new part numbers created during the pilot for which an approved equivalent already existed. This is the metric with the clearest path to a cost figure, since every avoided part number removes drawing, approval, tooling, and inventory work downstream.
Time to answer a standards or calculation question. Questions that currently route to one senior engineer or to an outside consultant.
Rework traceable to information that existed but was not found. Harder to instrument, and the most persuasive number you will produce if you can.
Weekly active use inside the cohort. A leading indicator rather than a result. If use decays after week three, the other four metrics will not save the pilot.
Baseline each of these before the cohort logs in, write the threshold next to it, and put a decision date in the calendar. A pilot with a decision date behaves differently from one without: the questions get sharper, and the work of instrumenting the baseline actually happens. For a fuller treatment of turning these into a defensible financial case, see our guide to measuring AI return for mechanical engineering teams.
Choose the workload, not just the tool
Teams tend to scope a pilot around a product and then look for things to do with it. Invert that. Choose one workload that is frequent, measurable, and genuinely irritating, then ask which system serves it.
For most mechanical engineering teams the best first workload is retrieval across everything the organization already owns. It is frequent, every engineer feels it, and it has an unambiguous baseline. It also has the useful property of being hard in exactly the way that separates capable systems from demonstrations: an assistant that answers well from a curated sample and poorly from a real vault will be exposed within days.
This is where Leo is designed to sit. Leo is an AI assistant for mechanical engineers, trained on more than one million pages of standards, textbooks, and technical articles, and connected to an organization's own knowledge base rather than to a generic corpus. Leo offers integrations with leading PDM and PLM platforms (SolidWorks PDM, Autodesk Vault, PTC Windchill, Siemens Teamcenter, Arena PLM, and others), along with local and network directories and ERP data, which is what makes a retrieval pilot testable against real history rather than a sample. Leo is SOC-2 certified and GDPR compliant, no AI is trained on customer data, and IP stays protected, which is usually the first question a pilot raises once real vault access is on the table.
Two decisions follow from the workload choice. The first is whether to buy a purpose built system or build retrieval in house, which we work through in buy versus build for AI on PLM data. The second is what to actually examine during evaluation, covered in what to look for in an AI copilot for mechanical engineers. If the pilot depends on vault access, the practical constraints in our note on AI for PDM and PLM integration are worth reading before week one rather than during it.
The 90 day loop, week by week
Ninety days is the right length. The NANDA research found mid-market organizations moving from pilot to implementation in roughly that window, while large enterprises took nine months or longer, and the shorter cycle is not a shortcut. It is what forces scope discipline.
A structure that holds up in practice:
Weeks 1 and 2, connect and baseline. Stand up access to real data, not a copy. Run the timed retrieval task set and record every number. Nobody in the cohort has access yet.
Weeks 3 and 4, onboard deliberately. Bring the cohort in together, in one session, with the workload named. Teams that hand out logins and hope get a decay curve instead of a result. Our guidance on accelerating engineering onboarding with AI knowledge management applies to the cohort itself, not only to new hires.
Weeks 5 to 8, run and observe. Weekly 20 minute conversations with each user. Log the questions the system answered badly, which is the most valuable artefact the pilot produces, because it tells you whether the failures are content gaps in your own vault or capability gaps in the tool.
Weeks 9 and 10, re-measure and stress the edges. Repeat the identical baseline task set. Then deliberately test the cases you avoided in week five: the legacy programme, the acquired product line, the supplier drawings nobody owns.
Weeks 11 and 12, decide and write it down. Compare each metric against its threshold. Produce a short memo with the numbers, the failure log, and a recommendation. A pilot that ends in a memo is reusable knowledge whichever way it lands.
One rule holds the whole loop together: do not expand scope mid-pilot. Every new workload added in week six resets the measurement and pushes the decision date out, and repeated cycles of that produce the organizational fatigue that makes the next pilot harder to run than this one.
FAQ
MIT Project NANDA, The GenAI Divide: State of AI in Business 2025, July 2025. Multi-method study conducted January to June 2025: a systematic review of more than 300 publicly disclosed AI initiatives, 52 structured interviews, and 153 senior-leader survey responses. Cited here for the evaluation-to-pilot-to-production funnel (60 percent, 20 percent, 5 percent), the finding that the root cause is a learning gap rather than model capability, and the mid-market pilot-to-implementation window of roughly 90 days against nine months or longer at large enterprises.
Run a pilot that ends in a decision
See how engineering teams scope, measure, and scale an AI pilot with Leo
Book a walkthrough. We will map your first pilot workload, agree the baseline metrics, and show how Leo connects to the PDM and PLM data your engineers already work in.
Schedule a Demo →
#1 New AI Software Globally - G2 2026
Enterprise-grade security
Trusted by world-class engineering teams
