ITADN

[Question]: dataset license

#11210Closedzhongww1 创建于 2026-01-27
questionstale
Z
zhongww1commented
### 请提出你的问题 Hi Team, Thank you for releasing the uie model. It has been very helpful for downstream Chinese NLP tasks. I would like to inquire about the training corpus used for pretraining uie, specifically regarding: The source(s) of the training corpus (e.g., public datasets, proprietary datasets, web-scraped corpora, or internal datasets). Whether the training corpus has identifiable dataset names that can be cited or referenced. The licensing terms associated with the corpus, especially in relation to commercial usage scenarios. Whether Paddle team can provide formal documentation or clarification on dataset licensing if required for compliance purposes. I noticed that the model itself is released under Apache-2.0 license on Github, but the training data licenses are not explicitly described. For users considering deployment in commercial or regulated environments, clarity on data provenance and licensing is important. Any additional documentation, links, dataset descriptions, or clarification would be greatly appreciated. Thank you in advance.
关闭于 2026-04-12 2 条评论