Business

Selecting an AI provider: procurement check for Austrian SMEs

Choosing AI providers: How Austrian SMEs compare benefits, test quality, data flow, contract, total costs, and exit before making a decision.

An adult business manager and an IT professional compare three neutral AI solution modules in an Austrian SME using colored test samples

A convincing AI demo is not yet a reliable purchase decision. In the prepared example the system recognizes documents, answers questions and generates reports in seconds. In actual operations, old file formats, Austrian technical terms, access-rights concepts, seasonal load and real edge cases meet the product. What seemed simple in the sales conversation becomes an integration and control task.

A structured AI provider check for SMEs separates benefit, risk and contract. Austrian companies do not need a months-long RFP project for this. They need a clear use case, realistic tests, comparable answers and an exit before data or processes are locked in. This guide leads from the requirements sheet to a justified procurement decision.

Procurement starts with the problem, not the model

First write on one page what isn’t working well today. Which process step takes time? Who is affected? What errors occur? What outcome should improve? An example would be: "Incoming maintenance reports should be pre-sorted by device type and urgency, without the system automatically releasing work orders itself."

This limitation prevents a small assistant in the sales conversation from becoming a platform for the entire company. Explicitly define what is not part of the assignment: no personnel evaluation, no automatic customer decision, no processing of special categories of data, no publication without approval.

Add three measurable success criteria and two stop criteria. Success could be a shorter turnaround time with an unchanged error rate. A stop criterion would be that sources are not traceable or data is stored contrary to the specification.

A comparable list of providers is generated from the requirements sheet

Divide the requirements into four groups: functional, technical, security and operational. Mark each requirement as mandatory, assessable or optional. A provider who does not meet a mandatory requirement is disqualified; many nice additional features must not obscure that.

  • Functional: supported documents, language, source display, correction capability and approval steps;
  • Technical: interfaces, identity management, roles, protocols, export and test environment;
  • Security: data locations, subcontractors, encryption, deletion, training and tenant separation;
  • Operations: Availability, support hours, changes, pricing model, exit strategy, and incident response time.

Send all vendors the same questionnaire and the same test cases. Answers like 'enterprise-grade', 'GDPR-compliant' or 'responsible AI' do not count as proof. Require concrete documents, configurations, responsibilities, and contractual clauses.

The data flow must be clear before pricing.

Map the path of an input: from the workstation through interfaces and providers to logs, support systems, backups and subcontractors. That data being 'hosted in Europe' does not automatically answer who can access it, where support takes place, or whether metadata goes to other services.

Ask precise questions:

  1. Which content, usage, and diagnostic data are processed?
  2. For what purposes are inputs and outputs stored?
  3. Are customer data used for training, improvement, or human review?
  4. Which subcontractors are involved, and how are changes announced?
  5. What retention and deletion periods apply to active systems, logs, and backups?
  6. How can access, rectification, deletion, and export be practically supported?

The WKO GDPR checklist reminds, among other things, of purpose, legal basis, processors, international data flows, security measures, and possible data protection impact assessment. These points do not belong in a later compliance round, but in product selection.

Check training, conversation history, and support access separately

A promise like "We do not train on your data" is important but incomplete. Content could still be stored for conversation history, abuse detection, error analysis, or support. Can individual features be disabled? Do different rules apply to administrative data? How long do deleted content remain in backups? Can support view content, and is that access logged?

Also check all product variants. A free web version, a corporate account, and an API may have different terms. The approval must refer to the specific plan, the specific region, and the activated features. Employees must not switch to a different account for the sake of convenience.

For internal preparation, the guide on "trade secrets in AI tools" is helpful. It defines which data classes are allowed to enter a test at all.

A realistic test corpus instead of a stage demo

Compile 20 to 50 representative test cases. They should include typical tasks, difficult exceptions, and deliberately unsolvable cases. Personal references and secrets are removed or replaced with synthetic data. For each case, experts define in advance which key statements, sources, or actions are expected.

The provider conducts the tests under the subsequent conditions: the same plan, the same rights, the same interface, and, if possible, the same model version. A prompt hand-optimized by sales in a different environment does not provide a reliable statement.

Do not just evaluate whether an answer sounds good. Measure technical accuracy, completeness, source reference, permissible refusal, stability upon repetition, and the effort required for human review. For integrations, rights checking, error handling, speed, and loggability are added.

Five tests every AI solution should pass

  1. Normal case: The system processes a common task completely and in a transparent manner.
  2. Edge case: Ambiguous information leads to a clarifying question instead of fabricated facts.
  3. Permission case: A person without permission is not given access to confidential sources or actions.
  4. Failure case: If unavailable, a manual route remains available and data will not be lost.
  5. Change case: After model or product updates, the most important tests can be repeated.

Save the input, configuration, expected result, and evaluation. This creates a small acceptance suite that can be reused later when changes occur.

Security is more than a certificate logo

Certifications and audit reports can build trust, but they do not automatically cover your own use case. Ask about scope, date, exceptions, and specific services. A certificate for the data center may say little about a subcontractor's new AI feature.

The ENISA provides SMEs with a procurement guide for cloud services with targeted security questions. When applied to AI, these include identity and access management, encryption, protocols, recovery, vulnerability management, and traceable security contacts.

Have the provider explain how they handle prompt injection, malicious files, data exfiltration and abusive actions. A system that only generates text requires different safeguards than an agent with access to email, CRM or accounting. Permissions are granted according to the principle of least privilege.

Translate contractual questions into clear operational obligations

The contract should include more than general liability and price. For the specific service, at least the following points are important:

  • Service description and the expressly permitted purpose of use;
  • Data protection roles and the required agreement on data processing;
  • Location of processing, subprocessors and change procedures;
  • Rules on training, secondary use, confidentiality and rights in inputs and outputs;
  • Availability, support, maintenance windows and response times;
  • Reporting of security, privacy, and serious AI incidents;
  • Changes to model, features, safeguards, and terms;
  • Export format, deletion, transitional support, and contract termination.

The European Commission publishes standard contractual clauses for controllers and processors. Whether and which clauses are appropriate must be examined by experts for the data flow; a download does not replace a contract review.

Limit the provider's rights to make changes

AI services are continuously changing. A different model, a new standard feature, or a changed subcontractor can shift quality and risk. The contract and internal processes must specify which changes will be announced, what information the customer receives, and when retesting is required.

Automatic activations are particularly critical. A new web search, a longer history, or an agent function must not enter an already approved process unnoticed. Administrators should check release information, configuration, and permissions. For material changes, there needs to be a blocking or rollback path.

The exit test should be performed before signing.

Ask how the company will leave the service. Which data, prompts, knowledge bases, logs, and configurations can be exported in a documented format? How long will the export be available after termination? When are active data and backups deleted? What confirmation is provided?

Test the export during the pilot, not only at the end of the contract. A PDF report is not a usable export if operations need to reuse structured knowledge entries or audit histories. Also maintain a manual fallback process until the dependency is consciously accepted and secured.

The NIST AI Risk Management FrameworkIt explicitly incorporates third‑party and supply‑chain risks and recommends contingency processes for outages or incidents. The voluntary framework does not provide Austrian legal advice, but it is a useful structuring tool for governance and auditing.

Compare total cost instead of the license price

A low price per user can become expensive if integration, testing, and operations are missing. Account for at least the following items over twelve to 24 months:

  • Licenses, usage‑based costs, and minimum purchase commitments;
  • Setup, data cleansing, interfaces and access control model;
  • Training, internal support and support packages;
  • Ongoing functional spot checks, security tests and documentation;
  • Remediation for errors as well as manual fallback;
  • Price adjustments, overage and exit costs.

Model three volumes: expected usage, double usage and low usage. Some pricing plans penalize growth, others make a small pilot disproportionately expensive. Ask about limits, throttling and cost controls before the team goes into production.

A weighted decision matrix without number tricks

Weight the criteria before the demos. For example: functional quality 30 percent, data protection and security 25, operations and integration 20, contract and exit 15, total cost 10. Mandatory criteria remain grounds for exclusion and cannot be offset by points.

Every evaluation requires a brief proof: a test result, contracting authority, technical documentation, or an explicit statement from the provider. Mark unknowns as unknown, not as an average score. Uncertainty is a risk in itself and must be clarified or consciously accepted before making a decision.

Have at minimum the business unit, IT, data-protection or legal, and the eventual users evaluate it. Procurement moderates comparability; it cannot determine technical suitability on its own.

Consult references with the same application profile.

A generic customer list says little about your own deployment. Request a conversation with an organization of similar size, data situation, and integration model. A corporation with its own AI team is not a meaningful reference for an SME that has to operate the service with two administrators.

Ask operational questions: How long did the rollout really take? What data cleansing was necessary? Which errors only appeared after the pilot? How does support respond to a critical issue? How often do the model, prices, or features change? Which tasks remain permanently in-house? Also ask what the reference would do differently today.

Treat references as supplementary evidence, not as a substitute for your own tests. Customers may use a different pricing tier, contract, or data location. Additionally, the provider will naturally select satisfied contacts. Record which statements are transferable to your case and which are not.

If no suitable reference is available, the importance of a reversible pilot increases. That is not an automatic exclusion for a young product, but a clearly documented uncertainty. It can be limited by shorter contract terms, smaller data volume, additional tests, or a lower initial budget.

The pilot remains reversible and limited.

A pilot uses a small number of trained accounts, a defined data class, and a controlled process. It has a start, an end, a budget, assigned owners, and a termination rule. Production decisions or external publications remain human-approved. The team documents errors and improvement ideas without copying unvetted data into private tools.

Use the guide for a secure AI deployment plan, to record roles, reviews and success criteria. The pilot is not automatically extended at its end. The decision is made based on the predefined criteria and the contract status as actually reviewed.

Twelve questions for the final vendor meeting

  1. Which specific task does your product solve in our chosen configuration?
  2. Which cases should it explicitly not handle?
  3. Which data are stored where, for how long, and for what purpose?
  4. Are content or usage data used for training or product improvement?
  5. Which subcontractors are involved and how are changes announced?
  6. How do roles, logging, deletion, and export work in practice?
  7. What quality limits and known error patterns do you document?
  8. How can we repeat our acceptance tests after updates?
  9. Who reports incidents within what timeframe to which contact?
  10. What availability and support response time are contractually guaranteed?
  11. What does the realistic volume and integration scenario cost over two years?
  12. How do we get our data back and prove deletion after the contract ends?

Good procurement also buys control

The best provider isn’t necessarily the one with the most well-known model or the most impressive demo. For an SME, what matters is whether the solution works reliably within its own processes, discloses limitations, supports security requirements, and can be operated without an unacceptable dependency.

A AI provider check for SMEs makes these differences visible: same test corpus, clear must-have criteria, verified data flow, a robust contract, realistic total costs, and a tried-and-tested exit. This makes the purchasing decision slower than a click, but significantly faster than the later repair of an unsuitable system.