Managers rarely need to write the best prompts themselves when it comes to artificial intelligence. Their more important task is to organize responsibility within the team so that benefits, quality, and protection interests align. Who may use an AI result? Who detects an error? Who stops a problematic process? And who explains a decision to employees, customers, or executive management?
A AI checklist for managers übersetzt diese Fragen in beobachtbare Arbeitsschritte. Sie hilft österreichischen Teamleitungen, nicht jeden Anwendungsfall neu zu erfinden und dennoch den Kontext zu berücksichtigen. Der folgende Leitfaden richtet sich an Führungskräfte, die bereits erste KI-Werkzeuge im Team sehen oder einen geregelten Einsatz vorbereiten. Er ersetzt keine rechtliche Einzelfallprüfung, schafft aber eine belastbare operative Routine.
Managerial responsibility begins before tool selection
The discussion often starts with a product name: Which assistant is the most capable? For management practice, a different order makes more sense. First comes the task, then the risk, and only afterwards the tool. An AI can speed up a draft, prepare a decision, or automate a process. These three roles have very different consequences.
The Austrian AI Service Center of the RTR erläutert für Betreiber von Hochrisiko-KI-Systemen unter anderem Pflichten zu Betriebsanleitungen, passenden Eingabedaten, Überwachung und menschlicher Aufsicht. Die konkrete rechtliche Einordnung hängt vom System und seinem Zweck ab. Für jede Führungskraft lässt sich daraus eine allgemeine Lehre ziehen: Menschliche Aufsicht braucht nicht nur einen Namen, sondern Kompetenz, Zeit und echte Befugnis.
Responsibility does not disappear because a provider operates the model. Anyone who adopts an AI suggestion into a workflow must be able to understand and control the substantive consequences.
The 90-second check for new AI ideas
Before a team tests a new tool, four short questions help. A single red answer does not automatically mean a ban, but it requires a more thorough review:
- People: Does the output affect access to employment, income, performance evaluation, training, or other important opportunities?
- Data: Is personal, confidential, protected, or unnecessary information being processed?
- External impact: Does the output leave the team, get published, or is it presented to customers as reliable?
- Automation: Could an incorrect result trigger an action without effective human oversight?
If all answers are clearly "no", a limited, data-minimizing test can often be evaluated quickly. If one or more "yes" answers are given, a specialist review is required and, depending on the case, data protection, information security, HR, legal counsel, or the works council may need to be involved.
The leadership checklist: twelve points for everyday work
1. State the purpose in one sentence
"We want to use AI" is not a purpose. Better is: "The team will create an initial draft of internal training questions from already approved product information." The sentence specifies the task, the data, and the target audience. If any of these points change, it will be re-evaluated.
2. Appoint a subject-matter responsible person
The person must be able to assess whether the output is professionally accurate. Experience with the tool alone is not sufficient. Different knowledge is required for a contract than for a social-media draft. The responsible person is explicitly granted the right to reject a result or stop the process.
3. Define permitted inputs concretely
A manager should use examples to define which information may be used. "No sensitive data" is too abstract. Lists are more meaningful: no application documents, no health data, no access credentials, no unpublished key figures, and no customer data in non-approved services. Where possible, texts are anonymized or reduced to the information that is truly necessary.
4. Name the worst plausible error
Teams check more carefully when they know what they are looking for. Can the AI state an incorrect deadline, discriminate against a person, disclose confidential data, or output a fabricated source? The worst plausible error determines the depth of the check.
5. Define the quality gate in advance
Before starting, it is decided which evidence a result requires. For texts, this can be original sources, four-eyes approval, and a fact log. For calculations, comparative values and random samples are required. For code, tests and security reviews are included. The WKO emphasizes in its Guideline on data and result quality, that AI content should be cross-checked with verified expert knowledge and trustworthy current sources.
6. Enable human oversight
The person responsible for oversight needs access to inputs, outputs, relevant sources and the system description. They also need time. A review scheduled only five minutes before publication is not effective oversight. For critical outputs, a second qualified person should be planned.
7. Involve affected parties and interfaces early
A team cannot decide alone when a system analyzes employee data or affects work performance. In Austria, questions of labor-law participation and information must be taken seriously. The Chamber of Labour provides information on AI in the workplace about rights and obligations concerning data protection, the AI Act and workplace co-determination. Managers should not wait to inform HR, data protection and the works council until after the contract has already been signed.
8. Create a secure testing environment
A pilot operates with a limited number of users, defined data and a fixed end date. Production decisions are not secretly shifted into the testing phase. For comparison, some tasks are also performed without AI. Only in this way will it become apparent whether time is actually saved or work is merely shifted to subsequent review.
9. Log errors and near-misses
It is not just major incidents that are instructive. A wrongly identified name, a distorted recommendation, or an inappropriate tone detected in time also show where controls are lacking. A short log with date, use case, effect, cause, and measure is enough to get started. It must not turn into a blame game for those who openly report a mistake.
10. Do not confuse results with human performance
AI can smooth out phrasing or standardize different ways of working. This must not lead to hasty performance evaluations. Especially in recruiting, target agreements, and personnel deployment, every manager needs a clear distinction between machine-generated suggestions, professional observation, and human decision-making. Application data deserves additional protection; the jobspot.at article on deleting application data explains important basics for handling data after a rejection.
11. Check the benefit with a few key figures
Time savings alone can be deceptive. A small set of metrics consisting of processing time, rework effort, error rate, acceptance, and complaints is useful. For external content, it can additionally be measured how often sources had to be corrected or statements had to be retracted. The benefit is weighed against license costs, training time, and control effort.
12. Prepare an exit decision
Every pilot needs criteria for pause or termination. Examples include recurring factual errors, uncontrollable data flows, lack of transparency, low acceptance, or a vendor change with unclear terms. A team that can only define success is not capable of making decisions.
A decision matrix for different consequences
Not every output needs the same approval level. A simple matrix links scope and potential impact:
- Low impact, internal: Draft or list of ideas; self-check by a trained female or trained male user.
- Medium impact, internal: Process proposal or summary of confidential material; expert review and only approved system.
- Medium impact, external: Kundenmail, Stellenanzeige oder Website-Text; fact and tone check, data verification and documented approval.
- High impact on people: Recruiting, deployment planning, evaluation or access to services; no automated decision, thorough review and involve the responsible internal units.
The matrix intentionally remains cautious. An apparently harmless summary can have high impact if it forms the basis of a personnel decision.
How leaders assess human oversight
The AI Act describes requirements for human oversight for high-risk systems. Regardless of the formal risk class, five practical tests help:
- Can the person overseeing the system explain what it is used for and where its limits are?
- Does the person see the relevant inputs and not just a finished result?
- Can they reasonably override, pause, or discard the output?
- Is their decision documented when significant consequences are possible?
- Is there a deputy so that oversight doesn't lapse during vacation or sick leave?
The WKO states in its current guideline"Humans have the final say in the use of AI"and recommends a four-eyes principle for critical outputs. Crucial is that this final say is professionally informed and not merely formally confirmed.
Address shadow AI without undermining trust
When employees use unauthorized tools, it's often because of a real work problem: time pressure, missing features, or an approval process that's too slow. A manager should first understand which task needed to be solved. After that, data risks and alternatives should be clarified. Blanket threats are more likely to drive usage underground.
Monthly office hours or a simple channel for new ideas are helpful. Anyone who reports a use case receives a response within a binding timeframe. At the same time, it remains clear: access credentials, personnel information, trade secrets and non-public customer data do not belong in arbitrary public tools.
The weekly 15-minute routine
Responsibility can be integrated into existing leadership work. A short weekly round with the active users is often enough to detect changes early:
- Which AI use cases were actually used this week?
- Were there incorrect, biased, or unexpected results?
- Were new types of data or external recipients involved?
- Has the tool, its configuration, or its contract noticeably changed?
- Does a team member need training, approval, or support?
Only deviations and decisions are logged. This creates a lean operational knowledge base without permanently storing every prompt detail.
Case study: AI drafts in customer service
An Austrian service team wants to prepare standard replies more quickly. Management defines the purpose as a drafting aid, not as automatic communication. The approved tool only receives anonymized case information. Names, contract numbers and special categories of personal data are removed. The output is compared with the knowledge base and approved by the responsible person.
In the pilot it becomes clear that processing time decreases, but outdated deadlines are occasionally phrased convincingly. The manager therefore adds a mandatory link to the current authoritative source and records the error rate in the weekly review. After four weeks the question is no longer whether “the AI is good” but whether the overall process with oversight works better. This perspective prevents an attractive individual output from being mistaken for a reliable workflow.
Warning signs that should prompt managers to pause immediately
- No one can explain which data the system receives or stores.
- An AI output influences personnel decisions without the affected individuals and the responsible parties being informed.
- Subject-matter oversight exists only on paper.
- The team must hide errors to achieve the promised time savings.
- A provider changes significant functions or terms without a new review.
- External content is published without checking sources or rights.
- Employees cannot stop or override a problematic output.
Frequently asked questions for team leaders
Does the manager have to review every AI result themselves?
No. However, the manager must ensure that a technically qualified person reviews it and that the level of control matches the risk. Responsibility is organized, not fulfilled by personally rubber-stamping every draft.
Is general AI training sufficient for the whole team?
A foundation creates shared understanding. In addition, exercises for the specific use case are needed: Which errors occur, which data are permitted, which sources apply, and when is escalation required? Article 4 of the AI Act focuses on knowledge, experience, and the context of use.
How detailed must the documentation be?
Detailed enough that purpose, responsible parties, approvals, and material deviations are traceable. Random wholesale storage of all inputs can create new data protection and security risks. Documentation should therefore be designed deliberately and purpose-specifically.
What is particularly important for AI in recruitment processes?
Personnel-related applications can have significant consequences for people. Leaders should not allow automatic selection, should ensure data minimization and human decision-making, and should involve HR, data protection and the works council early. The specific use case must be assessed technically and legally before it is launched.
Conclusion: Good leadership makes control visible
The most important leadership achievement in AI deployment is not a technical trick. It consists of neatly connecting purpose, data, decision, and control. A short checklist ensures that teams do not wait until after an incident to discuss responsibilities.
Start with an active use case. Perform the 90-second check, identify the worst plausible failure and appoint a technically authorized supervisor. Then agree on a quality gate, metrics, and abort criteria. This turns AI enthusiasm into a responsible work process.