Valuing an AI Asset in M&A: How It Differs from Valuing Software

Acquiring an AI company requires substantially more extensive due diligence than acquiring a conventional software developer. In addition to verifying title to the source code, an investor must assess rights in the model, training data, model weights and training infrastructure, as well as compliance with the EU AI Act, the GDPR and applicable licensing terms.

In this article, we examine the principal legal issues that must be reviewed before acquiring an AI business.

What are the key differences between an AI asset and conventional software? What should be considered when preparing a transaction involving an AI target, even where the transaction itself differs little in form from a standard acquisition of shares or a primary, cash-in investment?

Let us examine these issues.

This material will be useful to investors, funds, corporate purchasers, founders of AI start-ups, generative AI developers and the legal departments of technology companies.

What an AI Product Consists Of: Four Distinct Elements

From a legal perspective

In a conventional IT transaction, the intellectual property assessment focuses primarily on the product’s source code. Copyright subsists in the code and, depending on the jurisdiction, is either assigned to the employer or customer under an employment or commissioned development agreement, or vests in the company from the outset, as under the US work-made-for-hire model.

At the same time, the code may be protected as a trade secret and through properly documented confidentiality obligations binding employees and contractors. This protection is often supplemented by trade marks protecting the brand, licences covering the open-source components used in the product and patents protecting individual technical solutions. Taken together, these mechanisms define the scope of protection afforded to a conventional software product, which the investor examines during due diligence.

In the case of an AI product, however, it is not sufficient merely to verify whether the rights in the code have been transferred to the target or whether a trade mark has been registered. In practice, an AI product normally comprises four distinct elements, each of which is regulated and legally protected in a different manner.

Model architecture

This is a description of the structure of the neural network itself: how many layers it contains, which mathematical operations are used and how those operations are interconnected.

In practice, most architectures underpinning modern AI products are described in publicly available scientific publications. For example, the transformer architecture underpins the GPT and Claude families of language models, while convolutional neural networks form the basis of image recognition systems.

This means that the architecture itself does not generally represent the principal source of value. The assessment therefore shifts towards the following three elements.

Model weights

Model weights are the specific numerical values assigned to the parameters of a neural network as a result of its training.

While the architecture determines the model’s “form”, the weights record its “knowledge”, expressed through billions of numbers that determine how the model responds to a particular prompt. The principal economic value of an AI product is concentrated in these weights.

Data sets

These are the collections of data on which the model was trained, on which the quality of its performance was tested and on which it may subsequently be fine-tuned.

Both the model’s behaviour and the legal risks associated with its use depend directly on the quality and provenance of the data set.

Training infrastructure and pipeline

These are the internal tools, scripts and configurations that enable the model to be retrained or further developed without having to rebuild its architecture from scratch.

Without this element, the purchaser acquires only a snapshot of the model as it exists at a particular point in time, without acquiring the ability to develop it following completion of the transaction.

Stages of Model Development and Their Regulatory Features

Different risks arise, and different legal regimes apply, at each stage. Unless the relevant stage is identified, it is impossible to characterise accurately what the company has actually “done with the model” and, consequently, what it is transferring to the investor under the transaction.

Initial training stage: pre-training

At this stage, the model is trained “from scratch” on a large volume of data and acquires “general knowledge”, such as the ability to process language or recognise objects.

Pre-training is resource-intensive and, at the same time, presents the greatest risk in terms of the lawful use of data. As a rule, this stage is undertaken by the model developers themselves.

Fine-tuning stage

At this stage, an existing model is trained further for a particular task or industry vertical, for example using a body of medical reports or legal documents.

Fine-tuning most commonly becomes a matter of dispute in M&A transactions involving mid-market AI companies. Where the model could not have been fine-tuned without documents provided by a customer, the developer may no longer own the entire result of the fine-tuning outright.

Model application stage: inference

This is the operational stage at which the model processes user prompts.

No training takes place as such at this stage. However, separate risks arise, primarily in connection with the processing of data submitted by users and the extent to which the model’s output may reproduce third-party protected content.

Please note

There is an intermediate area between fine-tuning and inference: methods under which the model is not retrained but is instead “extended” by means of an external knowledge base.

One such approach is RAG, or Retrieval-Augmented Generation, under which the model accesses a connected database of the company’s documents when generating a response. Externally, this may appear to be “AI that knows your product”. Legally, however, it is closer to data processing than to the development of a proprietary model.

This distinction is fundamental when assessing a target. A company using RAG is not a model developer in the legal sense, and it would be inaccurate to describe it as an “AI developer” in the transaction documents.

Model Weights as a Separate Transaction Asset

The weights of a modern language model comprise a set of files containing billions of numerical parameters, with a total size ranging from hundreds of megabytes to hundreds of gigabytes.

Without the architecture, these numbers are useless in themselves. Nevertheless, they encode the entire useful result of the training process. Once the weights are known, the original product can be copied with relative ease.

Which forms of protection apply to model weights?

Neither the EU nor the United States currently has a bespoke legal regime specifically protecting the parameters of a trained model.

In practice, protection is achieved by combining existing legal mechanisms.

Copyright

In the United States, the US Copyright Office has consistently taken the position that legal protection requires a creative contribution by a human author.

Model parameters are, in substance, generated automatically during optimisation. Courts and regulators therefore tend to regard them as the result of a computational process rather than as a protected copyright work.

The position in the EU is less restrictive. It may theoretically be possible to classify model weights as a computer program or a database. However, there is currently no established case law on this issue, and it would be premature to rely on such a classification in a transaction.

Trade secrets

For closed-weight models, trade secret protection is currently the principal effective regime.

In both the United States and the EU, the definition of a trade secret covers information that has commercial value because it is not known to third parties, provided that its owner takes reasonable measures to preserve its confidentiality. By their nature, model parameters are capable of falling within this definition.

There is, however, a material limitation: trade secret protection is lost once the asset is made public. Open-weight models such as Meta’s Llama, Mistral and Falcon therefore lose the benefit of this regime, and the applicable licence terms effectively remain the only IP asset.

Contractual restrictions

Once trade secret protection has been lost or weakened through disclosure, for example when a model is deployed on a customer’s servers or when extended access is provided through an API, the use of the weights can be regulated contractually.

An agreement may impose express prohibitions on copying, reverse engineering where relevant, distillation, meaning the creation of a new model using the outputs of the original model, and disclosure to third parties. It may also impose field-of-use and use-case restrictions.

Where trade secret protection is formally weakened because customers gain access to the model, contractual terms compensate for a significant part of that weakening.

What does this mean for an investor?

Where a purchaser acquires an AI company, wording providing for the “transfer of all rights in the software” is insufficient in relation to model weights.

If a court subsequently refuses to recognise the weights as a “computer program” for the purposes of the applicable law, which is a material possibility in the United States, the purchaser may find that the transferred product is effectively protected only as a trade secret. The continued existence of that protection will, in turn, depend on whether the purchaser itself maintains the required level of confidentiality.

Recommendation: when preparing a transaction involving an AI target, include a separate technical schedule of the components to be transferred in the transaction documents.

For each component, the schedule should specify:

  • the applicable protection regime, such as trade secret protection, copyright, contractual restrictions or an open licence;
  • the storage location, including repositories, cloud storage and version checksums; and
  • the terms of the applicable open-source licences.

Such a schedule can simultaneously serve as an attachment to the seller’s representations and warranties and facilitate the practical transfer of the asset following completion.

Who Owns a Model Trained on Third-Party Data?

This is one of the most contentious issues in AI transactions in 2024–2026.

Almost every modern model has been trained using data for which the company did not have a direct licence. This may include materials obtained from publicly available internet sources, content distributed under public licences such as Creative Commons and, in some cases, data of questionable or plainly unlawful provenance.

The relevant questions are whether the use of such data for training itself constitutes an infringement and whether the associated risk is transferred to the purchaser together with the model.

The US Approach: Fair Use and Emerging Case Law

In the United States, the principal defence relied upon by AI companies is the doctrine of fair use. Under this doctrine, certain uses of protected material are permitted without the right holder’s consent where justified by the purpose and character of the use.

The application of this doctrine to the training of AI models is currently being tested in several landmark cases.

In The New York Times v. OpenAI & Microsoft, commenced in 2023, the newspaper brought claims against developers that had used its articles for training purposes. The claimant alleged that, in certain cases, the model reproduced extracts from its publications verbatim.

If the Supreme Court precedent in Andy Warhol Foundation for the Visual Arts, Inc. v. Goldsmith is applied in this case, training on articles without a licence may be found to constitute infringement.

The principle established by that precedent is that the use of a “copy” for the same commercial purpose as the original does not qualify as fair use and constitutes infringement. In OpenAI’s case, that purpose may be characterised as providing news information to users, which was also the newspaper’s original purpose.

OpenAI’s purpose of providing news information to users therefore coincides with the purpose of the newspaper, which also provides news. If the court finds that OpenAI should have purchased a licence, “free” training will cease to be regarded as fair use because it deprives right holders of their legitimate income from licensing training rights.

In Bartz v. Anthropic, in which an interim ruling was issued in June 2025, the court drew an important distinction for the market. Training a model using lawfully acquired books was characterised as fair use, whereas creating a training corpus using pirated copies was treated as infringement not protected by fair use.

Fair use is not a fact capable of being confirmed by a representation or warranty. It is a legal position that a company may advance if a claim is brought against it. An SPA should therefore not include a warranty stating that the training “constituted fair use”.

A warranty may, however, be given regarding the specific provenance of the training data and confirming that the company did not use clearly unlawful sources.

Recommendation: when acquiring or investing in an AI company, examine the provenance of each material category of training data in detail during due diligence and assess the categories of claims that could theoretically be brought in relation to it.

Any identified risks should be addressed through specific provisions in the SPA, including standalone or fundamental warranties and enhanced indemnity protection.

The European Union Approach: Data-Mining Exceptions and Regulatory Obligations

The European approach to training AI on third-party data is based less on case law and more on express statutory provisions.

The EU Directive on Copyright in the Digital Single Market, or DSM Directive, provides specific exceptions permitting text and data mining: the automated processing of texts and data in order to identify patterns.

For scientific purposes, this may be undertaken without restriction. For commercial purposes, it is permitted provided that the right holder has not opted out of such use in a machine-readable form, for example by means of specific notices on a website. Most European AI developers currently rely on this exception.

In September 2024, the first significant decision was issued by a German court in LAION v. Kneschke. The decision confirmed that non-commercial research use of data by the AI community may rely on this exception, whereas right holders’ opt-outs must be respected in the case of commercial use.

The EU Artificial Intelligence Act, or EU AI Act, imposes separate obligations on developers of large models. These include an obligation to disclose a general description of the data used for training and to maintain an internal copyright compliance policy.

Failure to comply with these requirements creates a separate regulatory risk that must now be taken into account when assessing an AI target.

An Additional Layer of Risk: Personal Data Used for Training

Where a training data set contains personal data, which is almost invariably the case for text models trained on internet-sourced data, the transaction is also subject to the GDPR, the European data protection regime.

The principal difficulty is that the right to erasure granted to individuals is difficult to exercise in relation to a model that has already been trained. A model cannot be made to “unlearn” specific data without substantial retraining.

European regulatory practice developed around this conflict in 2024–2025. In December 2024, the Italian data protection authority, the Garante, imposed a substantial fine of EUR 15 million on OpenAI on grounds connected specifically with the processing of personal data during training.

Please note: when preparing a transaction involving an AI target, compliance with data protection requirements in relation to training is frequently underestimated.

A standard data protection audit will generally review the handling of customer data within the product but may overlook the provenance of the training data.

Fine-Tuning a Model: Who May Claim Rights?

Fine-tuning is the stage at which the interests of the greatest number of parties intersect. It therefore requires the highest degree of contractual precision.

Where Company A takes an existing base model developed by Company B and fine-tunes it using its own data set, the result is, in substance, a new model because the base model now incorporates an additional layer of knowledge.

Depending on the structure of the relationships, several parties may claim rights in the resulting model:

  • the developer of the base model, under the terms of its licence;
  • Company A, as the party that carried out the fine-tuning, or “adaptation” in software copyright terminology;
  • the data set supplier, as the maker or compiler of the database, where the data set was not assembled by Company A itself; and
  • Company A’s customer, where the fine-tuning was carried out using the customer’s data and the agreement provides that rights in the resulting work belong to the customer.

In practice, this structure creates a complex combination of potential rights and obligations. Identifying and resolving them is a mandatory part of transaction preparation.

A typical risk: customer data used for training

A very common arrangement allows the model developer to “improve the model using aggregated data from all customers”, while the rights in those improvements remain with the developer.

This arrangement is currently under pressure from two directions: data protection law, because the legal basis for using such data for training may be questionable, and customers themselves, when they discover that their confidential information has effectively been used to develop a product made available to others.

When preparing the transaction, all customer contracts must be reviewed and the following specific question must be answered:

To what extent could customer data have been incorporated into the model’s training, and does the company have a clear contractual and regulatory basis for using the resulting improvements?

The absence of express customer consent, buried in standard terms of service, is a typical risk that may materialise after completion of the transaction in the event of action by a bad-faith competitor or an active regulator.

On-Premises vs SaaS: Different Transaction Risks

The method by which an AI product is distributed directly affects the allocation of risk in the transaction. The same product may be valued differently depending on how it is supplied.

A SaaS model, meaning Software as a Service, is a scenario in which the model weights remain physically under the developer’s control and the customer receives only online access to the model’s operation: the customer submits a prompt and receives a response.

The model parameters themselves are never transferred to the customer. OpenAI API, Anthropic API, Google Vertex and most modern AI start-ups operate under this model by default.

An on-premises model is a scenario in which the model is deployed within the customer’s own infrastructure. Under this model, the customer will usually obtain practical access to the model parameters and may use the model independently.

Many enterprise language models, secure medical models and industry-specific solutions in the financial and defence sectors operate under this model.

What does the customer acquire in each scenario?

In both cases, the principal asset remains the model parameters and the tools used to train it. However, the balance of risks differs.

Under the SaaS model, the model remains within the developer’s environment. The key risks therefore relate to compromise of the internal environment, including insider access and security incidents, and to the stability of the customer base.

Accordingly, when preparing a transaction involving such a target, particular attention must be paid to its information security posture, including:

  • certifications such as ISO 27001 and SOC 2;
  • access management systems;
  • activity and audit logs;
  • agreements with key customers; and
  • the terms governing the processing of customer data.

Under the on-premises model, the asset effectively leaves the company’s perimeter with each sale. The principal risk therefore shifts to licensing: a former customer holding a copy of the model may potentially become a competitor.

The terms of the licence agreements are consequently material. They should include express prohibitions on distillation and on transferring the model to third parties, together with technical protection measures such as binding the model to particular hardware.

Important: where a model is distributed on-premises, trade secret protection for the weights is effectively weakened with every new customer installation.

If an assessment of the target establishes that customer licences are poorly drafted, no technical protection measures are used and there is no contractual prohibition on distillation, the company’s assertion that its weights are protected as trade secrets becomes unreliable.

A Special Case: Products Built on Third-Party Base Models

A separate category of AI targets consists of companies whose products are built on the APIs of major providers such as OpenAI, Anthropic or Google.

From a technical perspective, such a product consists primarily of a set of model instructions, or prompt engineering, a connected customer knowledge base and a user interface.

From a legal perspective, the principal asset is not the model weights, which belong to the provider, but the agreement with the provider, the accumulated methodologies for working with the model and the user data.

When assessing such a company, it is critical to examine the provider’s terms governing the use of the base model.

For example, OpenAI’s business terms expressly prohibit the use of model outputs to train competing models and impose restrictions on resale.

The risk of dependence on a particular provider must also be assessed. To what extent can the product architecture be migrated to a different model if the provider changes its policies or pricing?

Licences for Open-Weight Models: Read the Terms Carefully

Where a target uses Llama, Mistral, Falcon or Stable Diffusion, the licence terms require separate and careful consideration. The outward appearance of open-source software may conceal material commercial restrictions.

Under the Llama licence, where the threshold of 700 million monthly active users is exceeded, the licence expressly prohibits “improving other large language models through distillation” and imposes mandatory attribution requirements.

Falcon, developed by TII, is distributed under a standard Apache 2.0 licence.

The licence applicable to a Mistral model depends on the particular version, ranging from Apache 2.0 for certain models to the Mistral Research License, which imposes restrictions on commercial use, for others.

Stable Diffusion is distributed under a licence containing express restrictions on particular use cases.

If the assessment establishes that the target is in breach of the licence terms applicable to its base model, this may have a material effect on the transaction. The right holder may terminate or withdraw the licence, effectively depriving the key asset of its commercial value.

Which Forms of Protection Actually Work?

Copyright: effective only for conventional software

The model architecture itself, meaning the mathematical description of its structure, is not protected by copyright in either the United States or the EU.

If the architecture has been published in a scientific paper, it becomes largely available to the public.

The source code of the training and operating infrastructure, including training scripts, internal administration tools and model lifecycle management tools or MLOps systems, is conventional software. Copyright protection applies to it in full, just as it would in an ordinary IT company.

Model weights remain a contentious area. Copyright protection is extremely limited in the United States. In the EU, it is theoretically possible, but this has not been confirmed in practice.

Data sets are protected at two levels.

Individual elements of a data set may be protected as separate copyright works. The company does not acquire copyright in third-party texts or images merely because they have been included in its training corpus.

The data set itself, as a composite object, may be protected as a compilation in the United States or as a database in the EU, provided that substantial investment has been made in its creation.

Model outputs are not protected by copyright in their unmodified form in the United States. Approaches vary between EU Member States but are moving in a similar direction.

Trade Secrets as the Principal Protection Regime and Their Vulnerability

As demonstrated above, copyright protection for the key components of an AI product is limited.

AI assets are therefore most commonly protected as trade secrets, particularly model weights and training infrastructure.

The advantages of this regime are that it:

  • does not require registration;
  • covers diverse elements, ranging from model weights to individual training parameters;
  • has no fixed term; and
  • operates in most jurisdictions.

A condition for maintaining trade secret protection is that the company must have taken “reasonable measures to maintain confidentiality”. This requirement determines the direction of the due diligence review when preparing a transaction.

An operational AI company should have:

  • a formal access-control system with activity logging;
  • confidentiality obligations in the employment contracts of employees who have access to the code and models;
  • non-disclosure agreements with all counterparties that have had access to key components;
  • storage systems marked as containing confidential information;
  • an established process for revoking access when employees leave and procedures for responding to incidents; and
  • a policy governing employees’ participation in independent open-source projects.

In the absence of these elements, the company’s assertion that its models are protected as trade secrets has no practical foundation, and the purchaser is acquiring an asset that is not legally protected.

Conclusion

There is currently no universal protection regime for an AI product. In practice, each company establishes a multi-layered system of protection in which each layer supports the others.

The architecture and general methods may be protected as trade secrets or know-how.

The source code may be protected by copyright and the terms of open-source licences.

The weights may be protected through trade secret protection and contractual terms. Where the product is distributed on-premises, these should be supplemented by technical protection measures such as encryption and hardware binding.

Data sets require a lawful basis for the use of the relevant data, including data-mining exceptions, licences or fair use, together with the specific database protection regime available in the EU.

Copyright, design patents and trade marks may be used to protect the user interface and branding.

Due diligence of AI companies is becoming an increasingly complex and comprehensive process.

A lawyer preparing such a transaction must therefore develop a thorough understanding of the product’s structure, verify that each layer of protection is effective in practice and avoid relying on a single protection regime or legal mechanism.

Where protection is weak at any level, for example where model weights are stored without encryption and are accessible to a broad group of employees, this becomes a separate reason to structure the transaction using additional warranties, enhanced indemnity protection and deferred consideration.

Authors: Inna Semenova and Yahor Kulazhenka.


This material was prepared for a third-party media platform: medium.com

How REVERA Can Help

Are you planning to acquire an AI company, raise investment or conduct a legal review of an AI product?

The REVERA team advises on international M&A transactions, investments in technology companies and comprehensive AI due diligence, including the assessment of intellectual property, training data, licences and compliance with the EU AI Act and the GDPR.

Where you are considering a transaction involving an AI business, our specialists will help identify legal risks before the transaction documents are signed and propose the optimal transaction structure.

 

Discuss AI Due Diligence

Request advice on an AI transaction