Codeberg Updates Terms of Use to Prohibit LLM Model Training Extrusions

1 min read
Codebergplatform Hacker Newspublisher

Codeberg's proposed terms of use amendment to prevent LLM training on its hosted code represents an important shift in how open-source platforms are addressing large-scale model training practices. For practitioners engaged in local LLM development and fine-tuning, this highlights the critical importance of understanding and respecting source licensing and platform policies when using public datasets and code repositories.

This development is particularly relevant for those building local models, whether fine-tuning existing architectures on specific codebases or curating training datasets from public sources. The push from platforms like Codeberg suggests the community expects explicit consent and attribution when code is used to train AI models, even locally-hosted ones. Practitioners should review the licenses of their training data sources and ensure compliance with repository terms of service.

The Codeberg policy discussion demonstrates broader industry momentum toward establishing clearer boundaries around data usage. For local LLM practitioners, this means being intentional about sourcing training data, properly licensing derivative models, and considering the ethical implications of where your training data originates.


Source: Hacker News · Relevance: 7/10