Companies

AI in Recruiting: Which HR decisions should not be automated

Using AI sensibly in recruiting: Which tasks HR can delegate, where human review remains necessary, and how Austrian teams can start fairly.

Three adult professionals in a Vienna meeting room jointly review an AI-assisted recruiting process

AI can buy recruiting teams time: it structures documents, identifies recurring requirements, and helps prepare for interviews. But it should not unobtrusively become the authority that filters people out. Especially where a tool evaluates applications, ranks them, or affects the visibility of vacancies, the issue is real access to the labor market. For Austrian employers, therefore, the decisive question is less whether a tool "uses AI" than which decision it actually prepares and who can be held transparently accountable for it.

This guide shows how HR can use an AI tool sensibly in recruiting: with clear boundaries, a genuine human decision, and a small verification process that also works for SMEs. It is aimed at people who are implementing or already using an applicant tracking system, a matching function, or a generative assistant.

Why AI in recruiting is more than a productivity tool

A text assistant that phrases an invitation more politely operates differently from a system that sorts applications by a point score. The first case supports communication. The second influences who is visible, who is invited to an interview, or who receives a rejection. That very effect is what counts.

The EU Commission explicitly names AI systems for recruiting and selection as a high-risk area when they, for example, analyze and filter applications or evaluate candidates. Targeted job advertisements can also be covered if the system substantially affects access to job opportunities. This is not a label for every spreadsheet helper. It is a reason to describe the specific use precisely before importing data or incorporating results into the selection process.

In practice this means: an AI result is an indication, not a personnel decision. Whoever only reads the top five profiles effectively adopts the system's preselection. A subsequent interview does not automatically remedy this mistake. It becomes particularly problematic when historical hiring data contain biased patterns or when the tool uses attributes as proxy signals that can be associated with gender, age, disability, origin, or caregiving responsibilities.

The first question: What may the tool do specifically?

Don't start with the provider's name; begin with a list of activities. For each function, write one sentence in the pattern: "The system receives X, produces Y, and Z uses the result for W." This turns a marketing claim into a verifiable use case.

  • Mostly supportive: Stellenanzeigen sprachlich überarbeiten, Interviewfragen aus einer freigegebenen Anforderungsliste vorbereiten, Termine koordinieren oder CVs in ein einheitliches Format übertragen.
  • Examine closely: Skills aus Lebensläufen ableiten, Profile matchen, Kandidat:innen in Kategorien einteilen oder Empfehlungen für die nächste Prozessstufe ausgeben.
  • Do not automate uncritically: Absagen auslösen, Rankings als verbindliche Shortlist verwenden, "suitability" aus Video, Stimme oder Online-Spuren ableiten oder Stellenanzeigen auf Personengruppen zuschneiden, die dadurch andere Jobs kaum sehen.

This classification also protects the team. A recruiter must be able to tell whether a score is merely a search aid or whether it turns a profile into an almost invisible hurdle. Therefore establish a simple traffic-light system: Green for administrative support, Yellow for recommendations with documented review, and Red for automated decisions with significant impact. Red is not released without legal, data protection, and, where applicable, works council review.

A human decision needs time, criteria, and a right to override

'Human in the loop' is only credible if the person in the process can actually decide differently. A click on 'confirm' after a long ranked list is not enough. Good human oversight has three components: the reviewing person knows the selection criteria, can view the underlying documents, and is allowed to overrule the result without disadvantage.

Example: For an assistant position, German language skills, proficient use of Office applications, and experience with customer contact are defined in advance as must-have or optional criteria. The tool may make documents searchable by these terms. However, it must not send an automatic rejection just because an equivalent experience was described differently. The recruiter reviews borderline cases using the same criteria, briefly documents their decision, and can correct the search result. This keeps the selection justifiable.

Also plan enough processing time. If a team receives 300 profiles per day, a 'complete manual review' without process design is not very credible. Instead, reduce the scope of use: for example, first improve data quality, make questions in the application form clearer, or filter by objectively required qualifications. AI must not hide poor process capacity.

The 30-minute check before the pilot

A pilot does not have to be bureaucratic. But it does need a robust start protocol. Choose a position that does not need to be filled urgently, and test with a small, as representative as possible selection of anonymized or lawfully used sample data. The goal is not to crown the 'best' AI, but to reveal failure patterns.

  1. Define the purpose: Which bottleneck does the tool solve? For example: detecting duplicates or structuring conversation notes. "Finding better applicants" is too vague.
  2. Note inputs and outputs: Which data goes in? What output does the team receive: summary, score, ranking or decision proposal?
  3. Publish selection criteria: Technical must-have criteria, nice-to-have criteria and exclusion reasons should be put into a short list before the test.
  4. Perform a cross-check: Have two people independently review some cases. Compare the tool's recommendation with their reasoning, not just with the final result.
  5. Search for error cases: Test different CV styles, career changers, part-time employment histories, longer gaps and qualifications from other countries. Especially here, missing keywords must not be read as missing competence.
  6. Define a stop rule: In cases of unexplained exclusions, conspicuous group differences, incorrect facts or data leaks, the pilot will be paused. Who decides on resumption?

Log only what is necessary: tool version, date, purpose, criteria tested, observed errors, responsible parties and the decision. This note helps later more than a long presentation because it shows why a deployment was approved, restricted or ended.

Privacy: Less data is often better data in recruiting

Application documents contain personal information, often also indirect indications of particularly sensitive life circumstances. Therefore, do not routinely upload them to publicly accessible chatbots. Clarify in advance where the data will be processed, whether the provider uses them for training, how long they will be stored and which contractual partners are involved. Even a tool that only "summarizes" can create a new processing step.

A data minimum proves practical: remove photo, date of birth, address, marital status and other information that is not required for the specific selection. Preferably do not use full CVs when a structured competency list suffices. Define roles: Who may export data? Who can see logs? When are test data deleted?

Involve data protection officers early. In Austrian companies, participation rights and information obligations toward the works council may also be relevant, especially when technical systems affect the behavior or performance of employees. Recruiting, IT, data protection and employee representation don't need to wait until after the contract has been signed to come together.

Detect bias without waiting for a magic metric

No single fairness metric proves that a selection process is fair. Still, you can specifically check whether a tool behaves suspiciously. For example, compare whether equivalent qualifications are treated similarly when phrased differently. Are career changers routinely devalued? Do candidates with longer employment gaps fail more often, even though that information is not decisive for the role? Are foreign languages, international degrees, or part-time work automatically interpreted as disadvantages?

The test must not itself create new sensitive profiles. Work with legally permissible, minimized test cases and discuss the method with data protection and equality officers. What matters is the process: an anomaly leads to a comprehensible adjustment, a more restricted use, or a stop. Anyone who only asks the vendor for a "bias-free" label shirks their own responsibility.

Contract and vendor: These questions should be on the table

Procurement is part of HR strategy. Ask the vendor for a clear description of the functionality, not just technical buzzwords. Particularly important for HR are data flows, change management, and support for complaints or corrections.

  • Which categories of data does the system process, and in which countries?
  • Are customer content or application data used for model training? Is that contractually excluded?
  • Which subcontractors are involved, and how are changes communicated?
  • Can the team export settings, criteria, and logs?
  • How does the provider explain a ranking or a recommendation in language recruiters can use?
  • Which security incidents, malfunctions, and significant model changes must be reported?
  • How does the provider help correct incorrect data or erroneous results?

A missing answer is itself a result. If a provider cannot explain either data usage or change logic, the product does not belong in a process that affects employment opportunities. Do not buy on a hunch, and do not rely on a demo with idealized profiles.

What the AI Act means for the roadmap now

Legal classification does not replace advice for individual cases. For planning, however, it is clear: systems in the area of employment can fall under the AI Act as high-risk. The European Commission currently indicates that the rules for certain high-risk areas, including employment, are expected to apply from 2 December 2027. Anyone selecting a recruiting system today should therefore design their architecture, data flows, and responsibilities so that later evidence can be provided.

This is not a reason to postpone useful assistance features. It is a reason to distinguish clearly. A scheduling tool needs different controls than a system that ranks candidates. Record each use case separately and do not silently change the purpose. A feature that is later switched from 'search' to 'automatic shortlisting' is not a small update but a new checkpoint.

A practical starting point for Austrian HR teams

Start small and where the benefit is easy to verify: for example, when quality-checking a job posting, summarizing interview notes from a closed system, or searching for already approved competencies. Measure not only minutes saved, but also correction effort, follow-up questions, and complaints. Set a fixed review date after four to six weeks.

Afterwards, the team decides based on the documented experience: keep the feature, narrow its boundaries, introduce additional controls, or discontinue its use. This loop doesn't make AI in recruiting perfect. But it prevents an apparently neutral recommendation from unintentionally determining applicants' career opportunities.

Implementing transparency in the application process

Applicants don't need a technical manual. However, they should be clearly informed when a relevant AI-supported process structures, evaluates, or prioritizes their documents. Phrase the information concretely: which step is supported? Which person reviews the result? How can affected individuals request a correction or ask questions? A general notice in a privacy policy does not replace this process clarity.

A short internal work instruction is also worthwhile. It should state the permitted purpose, prohibited inputs, the review steps for a score, and the contact for uncertainties. Especially important: recruiters must not present a recommendation as objective truth. In discussions with specialist departments it should remain clear that a ranking was produced only from the available data and settings. Professional suitability is assessed on the basis of the job requirements profile, not by a number on the screen.

A simple quality indicator is the number of subsequent corrections. If cases accumulate in which the tool overlooked skills, incorrectly summarized roles, or misclassified candidates, the team must not only rescue individual profiles. It investigates the cause: data format, criteria, configuration, or the entire use case. This feedback loop belongs in regular operations, not just in the first test month.

Conclusion: Recruiting remains an accountable decision

AI can relieve HR if it is used transparently, narrowly limited, and verifiably. The best rule is simple: people must not disappear because of a score that nobody can explain or correct. Those who clarify purpose, data, criteria, oversight, and stop rules before the pilot create a process that treats applicants more fairly and enables the company to make better long-term decisions.

So don't just test your next AI feature in the product demo. Use a realistic test case, have two people judge independently, and document what the system actually changes. This creates a reliable foundation for recruiting with AI in Austria.

Sources and further information