OverviewHistoryStatsSecurity
Security
Independent checks from skills.sh's audit partners.
Independent checks from skills.sh's audit partners.
Warn4 providerslatest audit Apr 16, 2026
Independent checks from skills.sh's audit partners.
This skill facilitates the benchmarking of code generation models using the BigCode Evaluation Harness. It allows for comprehensive evaluation across multiple languages and datasets. The primary security considerations involve the execution of model-generated code and the loading of remote code from model repositories. The skill acknowledges these risks and provides guidance on mitigation, such as using Docker for isolation.
Detected behaviors
No alerts
2 issues
Score: 93/100 · 2 sections analyzed