Bielik v3 11B Instruct
GPAI Provider: SpeakLeash
Metadata
Model URL: https://huggingface.co/speakleash/Bielik-11B-v3.0-Instruct
When was this model published? 2025-12-31
Is this a new or a fine-tuned model? Fine-tuned model
When was the Public Summary last updated? 2026-01-01
When did we evaluate the Public Summary? 2026-03-30 (see previous versions here)
Where did we find the Public Summary? We found this document linked on this page. An archived version can be found here.
Results of Evaluation
| Section | Transparency | Usefulness | General | |||||
|---|---|---|---|---|---|---|---|---|
| Clarity | Completeness | Consistency | Correctness | Accessibility | Comprehension | Transparency | Usefulness | |
| Document | 81.82 | 100.0 | 100.0 | 80.0 | 100.0 | 66.67 | 88.89 | 84.21 |
| General information | 80.68 | 86.59 | 100.0 | 100.0 | 100.0 | 90.91 | 87.58 | 93.75 |
| Public Data Sources | 100.0 | 83.87 | 100.0 | 100.0 | 100.0 | 100.0 | 92.28 | 100.0 |
| Private Data Sources | N/A | 100.0 | 100.0 | 100.0 | N/A | N/A | 100.0 | N/A |
| Scraped/Crawled Data | 91.6 | 81.48 | 100.0 | 77.88 | 0.0 | 67.95 | 83.78 | 51.46 |
| User Data | N/A | 100.0 | 100.0 | 100.0 | N/A | N/A | 100.0 | N/A |
| Synthetic/Other Data | 100.0 | 100.0 | 100.0 | 100.0 | N/A | 100.0 | 100.0 | 100.0 |
| Data Processing | 94.74 | 100.0 | 100.0 | 100.0 | 0.0 | 86.96 | 98.72 | 71.43 |
| Sum | 93.04 | 85.82 | 100.0 | 88.65 | 46.67 | 80.51 | 89.53 | 71.11 |
| Grades | B+ | C+ | ||||||
Evaluation Notes
If you update your public summary, please let us know so that we can update the evaluation and scores. We also welcome suggestions on improvements to this page, e.g. see FAQ section.
The open collaboration project SpeakLeash publishes a public summary for its Bielik v3 11B Instruct model on its website and also links it in its model card. The public summary follows the structure of the template, and describes Bielik as a fine-tuned version of Mistral. Of note, Bielik utilised the optional additional comments field in Section 1.3 to describe the purposes, regional, and linguistic characteristics of the dataset used as based on Polish and EU legal administrative domains and the inclusion of Silesian and Kashubian language Wikipedia pages. For their list of public data sources, however, the description referenced external documents ('preprints') which resulted in a deduction of points as relevant information was not provided in the document. We assessed its score to be 89.53% with grade B+ for transparency, and 71.11% with grade C+ for usefulness.
Suggested Improvements
The following are the suggested improvements based on using our our methodology where the public summary had issues related to the specified metric. The severity represents the extent of the issue, with low indicating aspects that could be fixable, and high representing missing information or requiring major changes.
- Document
- M1 Document should clearly indicate whether it is the latest version or if it is outdated and a replacement is made available.
Document should provide link to all versions of the document. low - M4 Document should provide link to authoritative source of the document. low
- M6 Document should clearly indicate changes from previous version, as well as where notice of updates or changes will be provided. medium
- M1 Document should clearly indicate whether it is the latest version or if it is outdated and a replacement is made available.
- General information
- M1 Date of data acquisition/collection for model training should have the proper format.
Languages should be described exactly (i.e. as a list of languages rather than as the number covered). low - M2 Links to additional publicly available documentation should be provided for the model(s).
The EU languages which are covered should be mentioned. low - M6 No indication of whether the model is continuously trained on new or dynamic data after the date provided. low
- M1 Date of data acquisition/collection for model training should have the proper format.
- Public Data Sources
- M2 Public dataset field does not only provide a list of large publicly available datasets, with explanations for selecting part of the datasets where necessary.
No information relating to the types of modality contained in the datasets. low
- M2 Public dataset field does not only provide a list of large publicly available datasets, with explanations for selecting part of the datasets where necessary.
- Scraped/Crawled Data
- M1 Field describing purpose of crawlers does not only contain information relating to the purposes of the crawlers listed. low
- M2 No list of the most relevant internet domains is provided, as per the requirements. low
- M4 No list of the most relevant internet domains is provided, as per the requirements. low
- M5 No list of the most relevant internet domains is provided, as per the requirements. high
- M6 No list of the most relevant internet domains is provided, as per the requirements. medium
- Data Processing
- M1 Information in 3.3 is not all relevant to data processing aspects and measures taken before or after training model. low
- M5 Not all information relating to the description of removal of illegal content is pertinent. high
- M6 Provider does not describe how they ensure that they are up to date with user rights requests. low
- Sum
- M1 Document should clearly indicate whether it is the latest version or if it is outdated and a replacement is made available.
Document should provide link to all versions of the document.Date of data acquisition/collection for model training should have the proper format.
Languages should be described exactly (i.e. as a list of languages rather than as the number covered).Field describing purpose of crawlers does not only contain information relating to the purposes of the crawlers listed.Information in 3.3 is not all relevant to data processing aspects and measures taken before or after training model. low - M2 Links to additional publicly available documentation should be provided for the model(s).
The EU languages which are covered should be mentioned.Public dataset field does not only provide a list of large publicly available datasets, with explanations for selecting part of the datasets where necessary.
No information relating to the types of modality contained in the datasets.No list of the most relevant internet domains is provided, as per the requirements. low - M4 Document should provide link to authoritative source of the document.No list of the most relevant internet domains is provided, as per the requirements. low
- M5 No list of the most relevant internet domains is provided, as per the requirements.Not all information relating to the description of removal of illegal content is pertinent. high
- M6 Document should clearly indicate changes from previous version, as well as where notice of updates or changes will be provided.No indication of whether the model is continuously trained on new or dynamic data after the date provided.No list of the most relevant internet domains is provided, as per the requirements.Provider does not describe how they ensure that they are up to date with user rights requests. low
- M1 Document should clearly indicate whether it is the latest version or if it is outdated and a replacement is made available.
Previous Versions
- 2026-03-30 (current version)
- 2026-01-12