Sep 25, 2026

How to make AI projects in manufacturing succeed: get the data right first

The AI projects that keep running after the demo share one thing: context under the data. Only 43% of the data manufacturers collect gets used.

How to make AI projects in manufacturing succeed: get the data right first

What the projects that keep running have in common

Of course you want your first AI project to still be running three months after the demo, answering the questions the morning meeting actually asks. That is possible on the data from the machines you have today, and the plants that get there do not have better models. They have better data underneath the model.

The difference only shows after the demo. A proof of concept runs on data from a spreadsheet someone exported by hand: clean timestamps, one machine, a column saying which order was running. Then it has to read live data. The project that survives that moment is the one where the order column also exists outside the export, where stop reasons are logged the same way on every line, and where the maintenance history carries a machine ID that matches the PLC.

None of that is AI work. It is data work, and it is the part you can start on this month.

A survey just measured how much room there is

Rockwell Automation's 2026 State of Smart Manufacturing report asked more than 1,500 manufacturers across 17 countries what is actually happening in their plants. One finding: of all the data these companies collect, only 43% is used effectively. The other 57% is already being generated and stored. It is waiting to be used.

The second number is the encouraging one. Asked to name the biggest internal obstacle to growth, 38% pointed at capturing, understanding, interpreting and using data. Budget constraints came second, at 36%, according to an analysis of the report. That ordering is new and worth noticing: the biggest lever is no longer what you can afford, it is what you can find out, and that lever is already inside your building.

Meanwhile a third of operations (34%) already run something AI-augmented, and the report expects more than half to be AI-supported by 2030. The plants that prepare their data now will be the ones for which that arrival is easy.

What usable data looks like on your floor

"Usable data" sounds abstract until you stand next to the line. In practice, a model can work with your data when this is true:

  • Every stop is logged with a start time and a duration, instead of a PLC counter that resets at midnight. You do not just know yesterday was bad; you can say which stop cost the most.
  • Tags have names an operator recognises. "Infeed conveyor, running" instead of DB12.DBW4, so the model and the planner who started in March read it the same way your senior operator does.
  • Downtime reasons come from the machine or from a fixed list at the moment of the stop, not from memory after the shift in free text like "jam", "infeed again", "same as Tuesday".
  • Inspection results are recorded, not just shown on a screen at the checkpoint.
  • The order number from ERP, the batch from the label and the machine data from the PLC are linked to each other, so it no longer takes a person who happened to be there to connect the three.
  • One definition of "running", so the same shift produces one OEE number and the morning meeting starts with the number instead of an argument about it.

The report says the same thing in survey language: connected, contextualised data builds trust, and trust is what turns information into action. That is a polite description of two people walking into a meeting with the same number.

Context is where the win sits

A value coming off a machine is not information yet. "21.4" is a number.

"Extruder 2, zone 3, 21.4 bar, 14:07, order 88213, batch F-2207, forty minutes into the changeover to the thinner film" is information. Everything after the 21.4 is context, and context is what makes a model useful.

This is the part nobody demos, because it is unglamorous: agreeing on names, joining machine events to orders and batches, deciding once what counts as a stop and what counts as a changeover. Do it properly and every later question gets cheap, because the join is already there. That is the real return on the work: not one AI project, but every question after it.

One shared data layer connecting PLCs, SCADA, MES and ERP so every machine event carries its order, batch and shift

That shared layer has a name in our trade, the unified namespace, and we explained it separately in what a unified namespace is and why AI needs one. What it buys you is simpler than the name suggests: one place where a machine event already knows which line, order, batch, product and shift it belongs to. MeshOS is how we build it, and the connections are read-only unless controlling something is the actual point of the project.

A number people can check is a number they use

The same report found that 93% of manufacturers expect to structurally reshape their workforce for smart manufacturing, and that 40% reskilled part of their workforce in the past twelve months. Training helps. What makes it stick is a dashboard the crew believes.

An operator believes a dashboard when it matches their day: when it shows 94% availability and they can click through to the two hours spent clearing the jammed infeed, listed as a stop with a time and a reason. So every number has to be clickable back to the events that produced it: this OEE, these stops, these timestamps, this is where the missing 6% went. A live OEE dashboard earns its place when you can drill from the percentage down to one stop at 03:12, and automatic downtime monitoring is what makes that drill possible.

At one packaging manufacturer we work with, every stop is pulled from the PLC and classified automatically: changeover, jam, starvation, microstop. The Pareto chart is not the achievement. The achievement is that nobody argues about it any more, so the meeting is about fixing the top stop.

The plan: one machine, then one line

  1. Pick the machine that hurts. Not the newest one. The one that shows up in every shift report.
  2. Read it where it stands. An edge computer next to the cabinet, read-only, on the PLCs and sensors that are already there. No machine gets replaced, no control logic changes.
  3. Name things once. A tag list an operator recognises. "Infeed conveyor, running" instead of DB12.DBW4, agreed once and used by everything that comes after.
  4. Add the context you already have. Order number, batch, product, shift. Most of it exists somewhere; it is just in a different system.
  5. Then ask your first AI question. By that point it is a small question, and you can check the answer yourself.

Two to four weeks per machine is a realistic pace. We do steps two and three for free on one machine, read-only, because looking at your own data is more useful than arguing about what it is worth: connect one machine.

Honest about what a data layer does not do

  • It does not produce a business case. Well-organised data that nobody acts on is the same problem in a more expensive form. Decide which decision the data is supposed to change before you connect anything.
  • Old controllers have a ceiling. A 1990s PLC may expose run, stop and a fault code and nothing else. Cycle-level detail there needs an added sensor, and that is an installation with a stop, not a software change.
  • Not every question needs AI. A threshold and an alarm solve more shop-floor problems than a model does, and they are far easier to explain when an auditor asks how the number was produced.
  • This survey is not your plant. More than half of the respondents were companies above one billion dollars in revenue. The direction is reliable, the benchmarks are not yours. A sixty-person plant with four lines has a different starting point, and usually an easier one.

If you want the longer version of the argument, including where industrial AI does and does not pay for itself, it is in our industrial AI whitepaper.

Frequently asked questions about AI on the factory floor

Do we need a unified namespace before we can start with AI?

No. You need context, and a unified namespace is one way to organise it. One machine, properly named and linked to orders and batches, is enough to answer a real question. Build the shared layer when the second and third line join, not as a precondition for the first.

How much data do we need?

It depends on the question, and the honest answer is usually less volume and more context. Anomaly detection on vibration can start with a few weeks of measurements. A model that predicts a specific failure needs examples of that failure, which can mean a year or more of history. A model that predicts rejects needs labelled rejects, so quality verdicts have to be recorded, not just displayed at the checkpoint.

Our machines are old. Can they be connected at all?

Usually yes. Siemens S7-300s, Modbus devices, dry contacts and retrofit sensors all work, and the machine keeps running while it happens. The limit is what the controller exposes. An older machine may give you running, stopped and a fault code. That is enough for availability and stop analysis. It is not enough for a quality model.

Can we keep the data in our own building?

Yes. MeshOS runs self-hosted on your own infrastructure as well as in the cloud, and the data is yours in both cases. For plants with automotive customers, or with an IT department that has NIS2 on its desk, self-hosting is usually the shorter conversation.

Conclusion: the successful project starts with measuring what you have

The room the survey measured is not a gap in ambition. Manufacturers are already collecting plenty of data; 57% of it is waiting to be used. That is why the projects that succeed start with the dull, cheap part: name the tags, join the machine events to orders and batches, make every number checkable, and only then ask a question worth modelling. Get that right and the AI part is the easy part.

Start with one machine and one question you already argue about every week. Once the answer holds up in the morning meeting, the second machine takes days instead of months.

Want to see what your own line produces? Connect one machine with us, or read how MeshOS brings production data together in one layer.

Back to Blog