Methodology

AI & Organizational Development

Designing the technology inside the work.

Introducing an AI system changes how people work, decide, learn and are assessed. Designing it is therefore an organisational problem before it is a technical one.

Whitepaper · 2026 · Alessandro Saccoia

01

Executive summary

An AI project is judged on the context it enters, the task it takes on, the effects it produces and who bears them.

An intelligent system enters a network of interdependent activities, roles, identities and meanings. Its organisational value depends on how it is embedded in the processes that hold that network together.

01 / Diagnose

Define the problem before the tool

A statement such as «people resist AI» is not a diagnosis. A diagnosis requires an explicit chain, from the construct to the dimensions, the indicators, the data, their interpretation and the way findings are given back.

02 / Configure

Automation and augmentation are configurations

They are not properties of the software but ways of organising work. The same system can widen a manager's room for judgement and narrow that of a team member.

03 / Develop

The tool is one part of the intervention

The intervention includes co-design, training, new practices, governance and evaluation. Without those components the organisation gets an installed piece of software and no change in how the work is done.

The argument

AI inside a company is a matter of organisational development before it is a matter of models. Accuracy is necessary but not sufficient, because the meaning of a system depends on how it redefines purposes, roles, coordination, rules, assessment and identity.

How to read this document

The nine pages follow four moves. You read the organisation and its culture, you diagnose with method, you develop the people, and you design and evaluate an AI system.

02

Where the technology enters

Six organisational processes to examine before designing a system.

An organisation is both formal structure and organizing, meaning the informal practices and tacit knowledge without which the real work does not function. A system built on the formal representation of the work always meets that margin.

01 / Purpose

Who defined the objective?

An algorithm optimised on a single indicator ends up sacrificing quality, learning or wellbeing. It is worth asking which dimensions were left out of the objective function.

02 / Differentiation

Different languages

Data science, HR and the front line mean different things by performance, error and fairness. AI puts specialisms in contact without a shared language.

03 / Integration

Coordination or a new divide

The system can coordinate information, or open a new divide between those who understand the model and those who merely follow its output.

04 / Formalisation

What is measurable and what counts

The algorithm translates criteria of judgement into computable rules. The risk is that what the system measures is taken for what has value.

05 / Assessment

Measurement changes behaviour

More frequency and precision can support learning, or produce surveillance, anxiety and opportunistic adaptation to the indicators.

06 / Identification

A threat to professional identity

When AI takes on a task that is central to a profession, the resistance is about the identity of whoever performed it rather than about ease of use.

03

Organisational culture

Beneath the visible artefacts sit espoused values and assumptions taken for granted.

Level 1 — visible

Artefacts

Spaces, technologies, dashboards, meetings, language, procedures. They show what happens, not why. A screen with real-time performance can express transparency and learning or control and competition, and the artefact alone does not say which.

Level 2 — espoused

Stated values

Innovation, autonomy, quality, inclusion. The distance between a stated value and the practice is diagnostic data. If whoever checks an output comes out less productive in the measurement systems, the value in practice is speed.

Level 3 — taken for granted

Basic assumptions

«Data is more reliable than people». «Only experience really understands this work». «An error is a fault to hide». They are written in no policy, yet they determine the reaction to AI. Where data is held to be superior the output becomes uncontestable, while where expert judgement counts the same system is felt as an attack on professional identity.

Depth, pervasiveness, stability

Culture is deep, pervasive and stable, and for that reason it gives meaning and is defended. It does not change because new values are proclaimed, but when structures, incentives, relationships and experiences change; if the new practices work for long enough they become credible and finally obvious.

Readiness belongs to a group, not to an organisation

Professional, generational and geographical subcultures read the same project differently. A system accepted by management can be refused by the professionals, and a system useful to experts can overload novices. For an assessment to be useful it has to be referred to a group, a use and precise conditions.

04

Diagnosis

An intervention fails when a vague problem becomes a solution too early.

01ConstructTrust, readiness, psychological safety
02DimensionsTheoretical decomposition of the concept
03IndicatorsWhat is observable and can be recorded
04QuestionsItems and guides, each with a reason
05DataSources triangulated across logs, voices and observation
06InterpretationAlternative hypotheses, not confirmations
07Giving backAlready part of the intervention, it opens dialogue or defence

«Acceptance» is not a single construct

Mandated use, actual use, perceived usefulness, trust and willingness to depend on the system are different constructs. A person can use it because they must and not trust it; find it useful and unfair; trust it on standard tasks and refuse it on career decisions. A high usage rate, where use is mandatory, does not demonstrate acceptance.

Divergence between sources is data

Management reads high adoption in the logs; operators describe ritual use and little trust. Triangulation serves to explain why the sources diverge, not to make them agree. An average of three can come from uniformly moderate answers or from half enthusiasts and half opponents, with opposite implications.

Tables and definitions in full in the whitepaper.

05

Developing people

Running a course, assigning a mentor or giving feedback is not yet development.

Training

Needs analysis before the course

If a person does not check the output because slowing down is punished, the problem is in the incentives. If the system is unreliable, the problem is the system and not how prepared the people are. Organisational analysis · Task analysis (KSA) · Person analysis An observable training objective asks someone to recognise four categories of risk in an output and to decide when to escalate.

Mentoring

Matching does not produce the relationship

Career functions give access and competence; psychosocial functions give identity and confidence. An algorithm that pairs two people does not «do mentoring». Sponsorship, coaching, exposure · Acceptance, role modelling · Initiation, cultivation, separation, redefinition Using private conversations to assess the relationship destroys the trust the programme set out to create.

Leadership

Leader development and leadership development

Leader development builds individual capability; leadership development builds the collective capacity for direction, alignment and commitment. Training many individuals does not produce the second. Identity and self-regulation · Deliberate practice · Assessment, challenge, support An app that tells the manager who to involve develops a leader; if only they see the data, power stays centralised.

Feedback does not always improve performance

It depends on where it directs attention. If it moves attention from the task to a self that feels threatened, it can make performance worse. Quality, timing, source credibility and the capacity for self-regulation count for more than frequency. An AI coach that increases frequency does not automatically increase learning.

Output is not outcome

«A hundred people trained» is an output. Reaction, learning, behaviour and results are not an automatic chain. A course people enjoy may teach nothing, and a competence acquired does not transfer if managers, tools and incentives do not support it.

06

Colleague or cage

The same technology becomes an algorithmic colleague or an algorithmic cage.

Algorithmic colleague

Judgement stays with the person

In a context that values judgement and autonomy the system supports without replacing responsibility. Divergences are discussed as a source of learning, the override remains practicable and tacit knowledge is cultivated.

Algorithmic cage

Autonomy erodes without a decision

In a hierarchical context it stiffens the processes and reduces autonomy. The override formally exists, but every deviation requires a justification, and agency erodes gradually without any decision ever having revoked it.

Five themes from the empirical research

  1. Human-AI collaboration produces benefits where there is task-technology fit, trust and the capability to use it.
  2. The algorithm is perceived as consistent and disinterested, or as decontextualised.
  3. Hope and fear coexist in the same person, and agency and the leader's support soften the fear.
  4. Algorithmic management assigns, monitors and sanctions, and the open question is contestability.
  5. Some technologies replace tasks, others create new ones.

Four dimensions of analysis

The manager is the first party using the system; the operator is also a second party, because the same dashboard measures them; the customer is a third party and bears the decision. Vendors and data annotators remain invisible actors.

How it evolves over time

Tables and definitions in full in the whitepaper.

07

Allocating the tasks

«Human + AI» on the same task is not always the best configuration.

Within-task complementarity justifies augmentation, because on the same task human and system together do better than either alone. Between-task complementarity instead justifies allocating tasks to whichever configuration suits each of them. In a study on an image classification task the two logics lead to appreciably different results.

68%Human only
77%AI only
80%Human with AI advice
88%Optimised allocation

Accuracy on an image classification task. The figures illustrate a design logic, not a benchmark transferable to other processes.

Easy cases / high confidence

Selective automation

With error audits and sampling. But if the simple cases disappear, new hires do not build basic competence.

Intermediate cases

Augmentation

The system orders the evidence, the person adds the context. It needs a real override, not a formal one, and explanations that can be used.

Hard cases / low confidence

Human teams

Multidisciplinary. Concentrating people only on the hard cases raises cognitive load and removes chances to recover and to learn.

Even a system on average less accurate than a person creates value when it is complementary or frees time for higher-value activity. Immediate performance does not close the decision, because responsibility, switching costs, meta-knowledge, fairness and the preservation of tacit knowledge weigh in the medium term.

08

Designing an AI system

Nine questions in order, from the context of the work to the risks of the system.

01ContextOrganisation, users, process, stakeholders, subcultures involved.
02NeedSpecific and supported by diagnosis, not deduced from an available technology.
03InputWhich data and knowledge, with what legitimacy and what minimisation.
04ProcessHow the system transforms the input and exactly where people intervene.
05OutputEvidence, alternatives and a confidence level, instead of a traffic light that hides the uncertainty.
06First stepA prototype on synthetic data and co-design, not ingestion of real conversations.
07Expected valueOn quality or resources, stated in advance and verifiable against a baseline.
08Human partsChoice of objectives, interpretation, relationship, decision and responsibility.
09RisksPrivacy, bias, drift towards assessment, dependence, misuse.

Tables and definitions in full in the whitepaper.

09

Evaluating in order to govern

Four families of KPI.

A useful evaluation plan holds four kinds of indicator together. Outputs say the system is running, mechanism indicators explain how it is used, outcomes measure the effects on the work, and risk indicators pick up the harms the first three families do not record.

Output

The system is running

People trained, messages generated, cases processed. Necessary but not sufficient.

Mechanism

How it is used

Cognitive load, comprehension, actual use of the advice, overrides and their outcome.

Outcome

Effects on the work

Behaviour at work, quality, timing, development outcomes, retention.

Risk

Harms the others do not record

Self-censorship, flattening of language, dependence, disparities between groups, incidents.

A trade-off is not a failure

If performance rises while psychological safety falls, the result is ambivalent and has to be decided on the size, distribution and duration of the effects. A single composite index would hide the tension.

A measure is valid for a use

A questionnaire useful for facilitating dialogue may be unfit for classifying people. If the system produces the metric by which it is judged, the indicator is not independent.