Francesco (@francedot) on X

· X (formerly Twitter) ·

10 min read Original article ↗

What the two attribution disputes have taught us regarding provenance, responsible reuse, and the sustainability of open-source research.

I had hoped not to have to write this article, but after trying to resolve two attribution disputes privately in the space of a month, I felt it was important to explain what happened and why it matters.

Cua Driver was developed so that using a computer would actually be useful for any coding agent, regardless of the harness, model, or vendor.

We selected the MIT License since we wanted researchers, developers, and companies to be able to build upon our work without any restrictions. This principle is still the same; Cua Driver will stay open source, and all the material that we have previously released under the MIT license will continue to be available on the same terms.

What open source needs is not just the right to reuse code but also the requirement to keep track of where that work originated.

The fact is that coding agents make the task more difficult since they are able to identify public implementations, translate code from one language to another, reproduce architectural patterns, and combine information from a number of different sources; the final code they produce might work perfectly and at the same time show no obvious trace of where it came from.

This presents a provenance problem for all the parties concerned: the maintainers, the developers, the companies, and the agents themselves.

Two recent examples

In the last month, two downstream AI implementations have cast doubt on the provenance of Cua Driver's native OS interaction work.

In the case of KimiCU, the MIT attribution provided by Cua Driver was initially omitted. We got in touch with the team on a private basis, and they admitted that the necessary notice had not been taken into account, before adding the correct copyright and license attribution in a subsequent release. We checked and confirmed that the attribution was included in KimiCU 0.5.8.

In that other instance, ZCode CUA’s implementation of the macOS background-click feature uses an unusually similar architectural combination to Cua Driver: accessibility hit-testing paired with a native, window-specific event path based on CoreGraphics and the private SkyLight APIs. Any one of these components might be reached independently; finding them assembled in the same feature is a more specific overlap. Given that specificity, we think it is reasonable to ask how the similarity arose, and we have asked. The distributed artifacts we examined did not mention Cua Driver, trycua, or its MIT license. Z.A.I. told us that the implementation was developed independently in 2022. We requested a dated commit, release, or design artifact supporting that timeline and have not received one.

Architectural similarity alone does not trigger the MIT notice requirement, and we are not claiming that the binary conclusively proves copying or a license violation. Our unresolved question is how this unusually similar implementation arose. If further evidence is provided, we will update this article.

Why this matters beyond two teams

It takes a great deal more than just writing the final code to keep an open-source project going.

We take the time to investigate undocumented behaviour of operating systems, design experiments that produce consistent results, check the findings on different platforms, record our observations, and publish the implementation so that other people need not go through the same research.

If the results of that work are used again without clearly stating where they came from, the cost will still have to be borne by the open-source maintainers even though the commercial and reputational advantages might go to others. The original authors may then have to look into downstream products, work out the sequence of development, and have to ask repeatedly for simple attribution.

That creates an unhealthy incentive structure:

  1. Teams that use open-source software carry out challenging research in public.
  2. Having obtained the results, other teams either reproduce them or reuse them.
  3. The implementation is included as part of a proprietary product.
  4. The original project will not receive any attribution, acknowledgement, or credit for collaboration.
  5. The people in charge take on the cost of working out what went wrong.

The issue isn't about reuse. In fact, we have selected the MIT licence because we do want people to reuse the Cua Driver.

The problem is using things without knowing where they came from.

There is an odd asymmetry here. In research, it is normal—even expected—to cite the papers and people whose work you build upon. A paper without references raises immediate questions about its provenance. Code is often treated as if the same standard does not apply: a team can build on a public implementation, package it into a product, and ship it without a visible reference to the project that made the work possible. The medium has changed, but the responsibility has not.

If it becomes the norm, there will be fewer teams who are willing to publish the most difficult and valuable sections of their work openly.

Why coding agents change the equation

If a team incorporated an open-source implementation, a human engineer typically being able to tell where it had come from, read the licence and pass that information on to the new project. Although that process was imperfect, there was still at least one person who could give an answer to the question: where did this code come from?

The process of agent-assisted development can end that chain. An agent might identify a public implementation, replicate its behaviour based on the documentation, translate a particular code path into another language, or combine information from a number of different projects. The developer who receives the output will see just a working result together with a brief explanation. The ironic part is that the result may still carry the same fingerprints: a few renamed variables or rearranged lines do not make a distinctive implementation independent, any more than changing the label on a package changes what is inside it.

The fact that there is uncertainty does not mean that the company is free from responsibility; if an agent contributes code to a commercial product then the organization which ships that product must understand the code's origin and keep any necessary notices.

It cannot be handed over to the model.

What responsible reuse should look like

Teams that are using coding agents should regard provenance as an engineering requirement, not merely as documentation completed at the end of a release.

At a minimum, agent-assisted development should:

  • record meaningful external sources consulted during implementation;
  • preserve copyright and license notices when code is copied or substantially adapted;
  • identify when an implementation has been translated, reconstructed, or behaviorally reproduced;
  • generate third-party notices for packaged applications and binaries, not just declared dependencies;
  • flag uncertainty when an agent cannot establish where a specific implementation came from; and
  • have to have human review carried out before the distribution of code which has unresolved provenance.

It is not necessary to have a perfect genealogy for every line of code; what is needed is a sincere effort in the case of an implementation that is particularly specific or obviously based on an existing project.

Companies should also ensure that it is easy for maintainers to raise a provenance concern without having to start a public argument. When responding in a responsible manner it is not necessary to accept each allegation at once; instead, they should look at the technical evidence, check the development history, and then give a specific answer.

In most cases all that is needed is to keep the upstream MIT notice.

What we are changing

We are by no means giving up on open source.

The core of Cua Driver will continue to be available to the public; the existing releases licensed under the MIT license will stay available on the same terms and nothing that has already been released is being taken back or placed behind a gate retroactively.

Experiences of this kind have led us to think again about the timing and manner in which we publish costly native research and new capabilities.

Going forward, we may:

  • delay publication of certain implementation details;
  • release some capabilities to trusted partners first;
  • gate experimental binaries or source code during early development;
  • require stronger provenance and attribution checks before distribution; and
  • maintain the core openness but make the decision to publish costly research as soon as possible.

It is not our intention to penalise developers who make a genuine mistake. The example of the KimiCU case illustrates the correct way to deal with a mistake: investigate the concern, include the necessary attribution, and then proceed.

The idea is that it might not be feasible to publish immediately without having a system in place to establish provenance.

The cultural shift we need

It is necessary for teams to instruct their agents regarding the responsibilities involved in using open-source software. An agent should not just focus on producing code that passes a test; it should also find out where the implementation originated from, identify the license that applies, and determine if attribution has to be included with the result.

Human engineers should look over those answers, in particular if the code involves undocumented APIs or if it replicates an unusual implementation pattern.

We need to teach agents to show the same care when dealing with the projects they develop as we do when working with human collaborators, such as by giving due credit, keeping a record of where information came from, and indicating uncertainty rather than quietly filling in the gaps.

Keeping open source open

Open source is not an extraction layer that has no provenance when it comes to commercial software; it is a system which is shared based on permission, attribution, trust, and reciprocity.

The MIT licence is designed to be permissive. It permits commercial use, modification, redistribution, and inclusion in proprietary products. The requirement to give attribution is intentionally kept low: when copying or redistributing substantial portions of the software, the copyright and licence notice must be preserved.

The relevant parts of the MIT License are straightforward:

Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the "Software"), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions:The above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software.

It is by no means an unreasonable burden.

We are keeping Cua Driver open source because we still believe in what open collaboration makes possible. At the same time, open source can remain sustainable only when the people doing the difficult work can publish it without losing the history, credit, and context that make collaboration possible.

Coding agents should make collaboration easier, not erase the path back to the people and projects that made the work possible.

If you want to help keep that work open, you can follow or contribute to Cua Driver on GitHub.

==

Update, September 2, 2026: We have clarified the ZCode section. Analysis of a distributed ZCode macOS build shows the same unusual combination Cua Driver uses for background clicks: accessibility hit-testing paired with a window-scoped CoreGraphics/SkyLight event path. That analysis cannot establish whether code was copied, translated, independently reconstructed, or developed earlier, and we do not claim that it can. Z.A.I. told us that its implementation was developed independently in 2022. We requested a dated artifact supporting that timeline and have not received one. If further evidence is provided, we will update this article.