Business

Trade secrets in AI tools: Which data should never be included in the prompt

Protect trade secrets in AI tools: How Austrian teams classify and minimize data before confidential content ends up in the prompt.

An adult Austrian information security specialist separates confidential documents from approved AI test data at an archival work table

An AI tool often needs only a few seconds to summarize a contract, compose a customer email, or explain source code. That very convenience tempts people to transmit more data than necessary. Names, internal prices, unpublished projects, or even access keys can end up in the prompt. Once input has been sent to an external service, it can't be retracted like a local typing error.

Austrian companies therefore need an easy-to-understand protection process for trade secrets in AI tools. The focus is not the question "Is AI allowed?" but rather: which information may be sent to which system for which purpose? This guide shows how teams classify data before input, minimize it, and respond appropriately in case of an error.

The prompt is a data transmission, not a scratchpad

Anyone who copies text into an external AI service transmits it to a provider. Depending on the product, contract and configuration, storage, logging, support access, model improvement, or international data flows may be regulated differently. A private free account is therefore not equivalent to a vetted corporate solution.

The WKO states in its current guideline on company-related data and confidentiality states that confidential information about the company itself or third parties should not be used in AI applications. API keys should be treated like passwords. It recommends clear internal rules, secure test environments and the integration of data protection.

An output can also be confidential. If an internal system searches sensitive sources and generates an answer from them, that answer must not automatically become public. Input, intermediate steps and the result need the same level of protection.

Four data classes for quick decisions

A practical classification is better than the blanket statement 'No sensitive data'. Four classes are sufficient for many companies:

Public

Information that the company has already consciously made publicly available, such as released website texts, published product sheets or public job postings. Here too, copyright, timeliness and purpose must be taken into account. 'Findable on the internet' does not automatically mean 'free to use'.

Internal

Work-related information without particular confidentiality that is nonetheless not intended for the public: internal templates, general process descriptions, or non-personal meeting notes. They may only be entered into explicitly approved systems and use cases.

Confidential

Trade and business secrets, unpublished prices, proposals, financial figures, strategies, customer details, source code, security architecture, contract contents, or information from partners. This category is generally excluded from general AI services. Use requires a verified technical and contractual protection pathway and explicit authorization.

Highly protected

Passwords, API keys, private keys, health data, personnel files, application documents, bank details, and other information with a high potential for misuse or harm. Such data does not belong in freely usable prompt fields. For a professionally necessary use of AI, a specially evaluated process is required.

Each category should be illustrated with examples from one’s own organization. A hotel has different secrets than a mechanical engineering firm, a tax consultancy, or a software company.

This information is especially easy to overlook

Data leaks do not only arise from complete documents. Often confidential details are copied inadvertently:

  • Email signatures with names, phone numbers, and direct contacts,
  • Comment columns and change-tracking in Office files,
  • Filenames with customer or project names,
  • Worksheets and hidden columns,
  • Metadata in images and documents,
  • Error messages containing server paths or tokens,
  • Source code with hard-coded credentials,
  • Long chat histories with earlier confidential contexts.

When copying from a CRM or spreadsheet, the clipboard may contain more than the visible selection. Teams should first check and clean data in a secure intermediate view, rather than pasting directly from the source system into the chat.

Anonymization is more than just replacing names

“Anna Huber” becomes “Customer A” – but combined with location, rare profession, contract date, and complaint details, the person may still be identifiable. This is pseudonymization, not automatically anonymization. Pseudonymized data remains personal data if identification is possible with additional knowledge.

The WKO points out in its notes on customer-related data regarding data minimization, anonymization, pseudonymization, and potential international data transfers. The European Data Protection Board addresses in its Opinion 28/2024 among other things, when AI models can be considered anonymous and which data protection issues arise during development and deployment.

In practice, the omission test helps: Can the task still be sensibly completed without this detail? If so, it is removed. For wording assistance, the type of request is usually sufficient; name, contract number, and exact address are superfluous.

The 15-second check before every input

A quick check at the workstation prevents many errors:

  1. Source: Who owns the information, and may I use it for this purpose?
  2. Content: Does it contain personal data, secrets, access credentials, or third-party rights?
  3. Necessity: What is the smallest portion necessary for the task?
  4. Tool: Is this exact system approved for this data class?
  5. Disclosure: Who is allowed to view, store, or share the result?

If a question cannot be answered with certainty, the information is excluded. The team uses a public sample file, synthetic data, or consults the responsible authority. Time pressure is not an approval mechanism.

Approved tools by purpose instead of brand name

A whitelist should not just say "Tool X allowed." It should specify user group, purpose, data class, account type, and key configuration. A service can be approved for public text ideas and blocked for customer documents. An enterprise instance can be evaluated differently than a private account of the same provider.

Relevant factors in the assessment include, among others:

  • are inputs or outputs used for training,
  • how long are contents stored,
  • where are data processed and which subcontractors are involved,
  • can logging and history be controlled,
  • are there roles, single sign-on and access restrictions,
  • how are data deleted or exported,
  • what happens at the end of the contract,
  • how does the provider report security incidents and changes?

The answers belong in the tool registry. Marketing claims like "Your data is safe" do not replace a concrete contractual and configuration review.

Take confidentiality toward customers and partners into account

A company must not only protect its own secrets. Offers, specifications, prototypes or analyses can be protected by contract or simply because of their clearly confidential nature. The fact that an employee has access to a file does not mean they are allowed to transmit it to another service provider.

When reviewing order documents, verify whether the agreed purpose includes processing by the chosen AI system. Relevant questions are: Is a confidentiality agreement in effect? May a subcontractor be involved? Has a specific storage location been promised? Are there deletion deadlines or requirements for returning documents? If unclear, the original remains in the customer's approved system.

Seemingly harmless excerpts can together form a secret. Individual technical measurements, delivery dates and location details may allow conclusions about an as-yet unpublished project. Therefore, the release should consider the entire context, not just each copied paragraph in isolation.

Carefully inspect source code and error messages

Developers often use AI for debugging. However, a stack trace can contain internal hostnames, file paths, user data or tokens. Source code can reveal proprietary logic, customer customizations and comments with confidential context. Before inputting, a minimal reproducible example is prepared.

To do this, the team removes access credentials, replaces real domains and IDs, truncates the code to the affected function, and checks license and customer bindings. A local static analysis tool or an approved company instance may be more appropriate for real code than a public chat. The response is then tested within the team's development process; it must not be deployed to production without review.

It is particularly important to separate secrets from error descriptions. Instead of sending an entire configuration file, it is often sufficient to provide: the technology used, expected behavior, an anonymized error message, and a small synthetic example. That way the model gets the technical structure, not the production keyring.

Secure alternatives to the full document

Often the benefit can be preserved without transferring the original:

  • use a structure with placeholders instead of the real contract,
  • check only a non-sensitive paragraph instead of the entire file,
  • generate sample data with a similar format but fabricated content,
  • use local search or editorial tools in a controlled environment,
  • use a vetted internal system with limited sources and permissions,
  • Transfer only categories or statistical values instead of individual cases.

Internal knowledge access must not become blanket full access. Even with retrieval-augmented generation or AI agents in the officeSources, users, and tools need minimal permissions. The system should only be able to find those documents that the requesting person is allowed to see.

Access keys must never be included in the prompt

API keys, passwords, tokens, and private keys can be immediately abused. They are stored in a designated secret store and provided to the technical process only at runtime. Neither a prompt template nor logs, screenshots, tickets, or source code files are appropriate storage locations.

If a key was accidentally transmitted, deleting it from the chat is not enough. It must be considered compromised: stop using it, revoke or rotate the key, check access logs, and report the incident according to the internal process. The absence of visible misuse does not prove that the key remained secret.

What to do if confidential data has already been entered?

Quick, orderly action reduces the damage:

  1. Do not enter any further content into the same conversation.
  2. Record the nature, scope, time, account, and affected system.
  3. Notify the internal contact for data protection or information security.
  4. Immediately lock or rotate access credentials.
  5. Use the provider's available deletion and support channels.
  6. Review affected processes, logs, and data disclosures.
  7. Evaluate the necessary legal reporting and notification steps with appropriate expertise.

Employees should be able to report a mistake without fear of pressure to cover it up. A secret kept too long is usually riskier than reporting an accidental misclick early.

A practical data filter for teams

A short form or preliminary editorial step can improve inputs. It asks for purpose and data class, flags patterns such as email addresses or keys, and requires approval for confidential classes. However, an automatic filter is no substitute for human classification. It may recognize a credit card number but not necessarily an undisclosed pricing strategy.

Test the filter with your own examples: Austrian phone numbers, company register data, internal project names, code snippets, and document comments. False-negative hits are documented just like unnecessary blocks.

Practical example: comparing proposals without customer secrets

A sales representative wants to compare three proposals with AI. The originals contain customer names, individual discounts, delivery terms, and internal contribution margins. Instead of uploading them in full, the team creates a shared comparison table with neutral vendor identifiers and only those criteria that are necessary for the decision.

The AI tool suggests a structure and open questions. The actual pricing decision remains with the responsible people who work with the original documents in the protected system. The AI history contains neither customer identity nor internal margin. The benefit — faster structured comparison — remains, while the risk to confidentiality is significantly reduced.

Checklist for a safe prompt

  • Is the account and tool approved for this purpose?
  • Has the information been assigned to a business data class?
  • Have names, contact details and unique identifiers been removed?
  • Does the input not contain any passwords, tokens, or API keys?
  • Have comments, metadata and hidden content been checked?
  • Is a snippet or a synthetic example sufficient?
  • Is it clear where output and history are stored?
  • Is the recipient allowed to view the result?
  • Is there a reachable approval authority in case of uncertainty?

Frequently asked questions about secrets and AI

Are paid business accounts automatically secure?

No. They may provide better protection and management features, but they must be specifically reviewed and properly configured. Crucial factors are the contract, data flow, account type, access, and purpose.

Can publicly discoverable content always be used?

Not automatically. Copyright, personality, trademark, and data protection rights may still apply. In addition, combined public information can enable new inferences.

Is pseudonymization sufficient?

It reduces risks but does not necessarily make data anonymous. If individuals remain identifiable with additional knowledge, data protection requirements continue to apply. Data minimization and the legal basis must be examined separately.

How do these rules fit into everyday work?

Embed data classes, a positive list and a reporting channel in the company's AI policy. Short examples and exercises are more effective than an abstract list of prohibitions.

Conclusion: Reduce first, then transfer

The most secure confidential dataset is the one that an external AI service never receives. Teams should therefore check the purpose, data class, necessity and tool before every input. Often an anonymized example or a small public excerpt is sufficient.

Start with five typical documents from your area and mark which fields are public, internal, confidential, or specially protected. This creates a concrete quick guide. If employees can recognize within a few seconds what must be kept out, data protection becomes part of the workflow instead of an after-the-fact fix.

Sources and further information