{"id":"bigcode-evaluation-harness","name":"evaluating-code-models","summary":"HumanEval、MBPP、MultiPL-E、15+ベンチマークのコード生成モデルをpass@k指標で評価します。","body":"# BigCode Evaluation Harness - Code Model Benchmarking\n\n## Quick Start\n\nBigCode Evaluation Harness evaluates code generation models across 15+ benchmarks including HumanEval, MBPP, and MultiPL-E (18 languages).\n\n**Installation**:\n```bash\ngit clone https://github.com/bigcode-project/bigcode-evaluation-harness.git\ncd bigcode-evaluation-harness\npip install -e .\naccelerate config\n```\n\n**Evaluate on HumanEval**:\n```bash\naccelerate launch main.py \\\n  --model bigcode/starcoder2-7b \\\n  --tasks humaneval \\\n  --max_length_generation 512 \\\n  --temperature 0.2 \\\n  --n_samples 20 \\\n  --batch_size 10 \\\n  --allow_code_execution \\\n  --save_generations\n```\n\n**View available tasks**:\n```bash\npython -c \"from bigcode_eval.tasks import ALL_TASKS; print(ALL_TASKS)\"\n```\n\n## Common Workflows\n\n### Workflow 1: Standard Code Benchmark Evaluation\n\nEvaluate model on core code benchmarks (HumanEval, MBPP, HumanEval+).\n\n**Checklist**:\n```\nCode Benchmark Evaluation:\n- [ ] Step 1: Choose benchmark suite\n- [ ] Step 2: Configure model and generation\n- [ ] Step 3: Run evaluation with code execution\n- [ ] Step 4: Analyze pass@k results\n```\n\n**Step 1: Choose benchmark suite**\n\n**Python code generation** (most common):\n- **HumanEval**: 164 handwritten problems, function completion\n- **HumanEval+**: Same 164 problems with 80× more tests (stricter)\n- **MBPP**: 500 crowd-sourced problems, entry-level difficulty\n- **MBPP+**: 399 curated problems with 35× more tests\n\n**Multi-language** (18 languages):\n- **MultiPL-E**: HumanEval/MBPP translated to C++, Java, JavaScript, Go, Rust, etc.\n\n**Advanced**:\n- **APPS**: 10,000 problems (introductory/interview/competition)\n- **DS-1000**: 1,000 data science problems across 7 libraries\n\n**Step 2: Configure model and generation**\n\n```bash\n# Standard HuggingFace model\naccelerate launch main.py \\\n  --model bigcode/starcoder2-7b \\\n  --tasks humaneval \\\n  --max_length_generation 512 \\\n  --temperature 0.2 \\\n  --do_sample True \\\n  --n_samples 200 \\\n  --batch_size 50 \\\n  --allow_code_execution\n\n# Quantized model (4-bit)\naccelerate launch main.py \\\n  --model codellama/CodeLlama-34b-hf \\\n  --tasks humaneval \\\n  --load_in_4bit \\\n  --max_length_generation 512 \\\n  --allow_code_execution\n\n# Custom/private model\naccelerate launch main.py \\\n  --model /path/to/my-code-model \\\n  --tasks humaneval \\\n  --trust_remote_code \\\n  --use_auth_token \\\n  --allow_code_execution\n```\n\n**Step 3: Run evaluation**\n\n```bash\n# Full evaluation with pass@k estimation (k=1,10,100)\naccelerate launch main.py \\\n  --model bigcode/starcoder2-7b \\\n  --tasks humaneval \\\n  --temperature 0.8 \\\n  --n_samples 200 \\\n  --batch_size 50 \\\n  --allow_code_execution \\\n  --save_generations \\\n  --metric_output_path results/starcoder2-humaneval.json\n```\n\n**Step 4: Analyze results**\n\nResults in `results/starcoder2-humaneval.json`:\n```json\n{\n  \"humaneval\": {\n    \"pass@1\": 0.354,\n    \"pass@10\": 0.521,\n    \"pass@100\": 0.689\n  },\n  \"config\": {\n    \"model\": \"bigcode/starcoder2-7b\",\n    \"temperature\": 0.8,\n    \"n_samples\": 200\n  }\n}\n```\n\n### Workflow 2: Multi-Language Evaluation (MultiPL-E)\n\nEvaluate code generation across 18 programming languages.\n\n**Checklist**:\n```\nMulti-Language Evaluation:\n- [ ] Step 1: Generate solutions (host machine)\n- [ ] Step 2: Run evaluation in Docker (safe execution)\n- [ ] Step 3: Compare across languages\n```\n\n**Step 1: Generate solutions on host**\n\n```bash\n# Generate without execution (safe)\naccelerate launch main.py \\\n  --model bigcode/starcoder2-7b \\\n  --tasks multiple-py,multiple-js,multiple-java,multiple-cpp \\\n  --max_length_generation 650 \\\n  --temperature 0.8 \\\n  --n_samples 50 \\\n  --batch_size 50 \\\n  --generation_only \\\n  --save_generations \\\n  --save_generations_path generations_multi.json\n```\n\n**Step 2: Evaluate in Docker container**\n\n```bash\n# Pull the MultiPL-E Docker image\ndocker pull ghcr.io/bigcode-project/evaluation-harness-multiple\n\n# Run evaluation inside container\ndocker run -v $(pwd)/generations_multi.json:/app/generations.json:ro \\\n  -it evaluation-harness-multiple python3 main.py \\\n  --model bigcode/starcoder2-7b \\\n  --tasks multiple-py,multiple-js,multiple-java,multiple-cpp \\\n  --load_generations_path /app/generations.json \\\n  --allow_code_execution \\\n  --n_samples 50\n```\n\n**Supported languages**: Python, JavaScript, Java, C++, Go, Rust, TypeScript, C#, PHP, Ruby, Swift, Kotlin, Scala, Perl, Julia, Lua, R, Racket\n\n### Workflow 3: Instruction-Tuned Model Evaluation\n\nEvaluate chat/instruction models with proper formatting.\n\n**Checklist**:\n```\nInstruction Model Evaluation:\n- [ ] Step 1: Use instruction-tuned tasks\n- [ ] Step 2: Configure instruction tokens\n- [ ] Step 3: Run evaluation\n```\n\n**Step 1: Choose instruction tasks**\n\n- **instruct-humaneval**: HumanEval with instruction prompts\n- **humanevalsynthesize-{lang}**: HumanEvalPack synthesis tasks\n\n**Step 2: Configure instruction tokens**\n\n```bash\n# For models with chat templates (e.g., CodeLlama-Instruct)\naccelerate launch main.py \\\n  --model codellama/CodeLlama-7b-Instruct-hf \\\n  --tasks instruct-humaneval \\\n  --instruction_tokens \"<s>[INST],</s>,[/INST]\" \\\n  --max_length_generation 512 \\\n  --allow_code_execution\n```\n\n**Step 3: HumanEvalPack for instruction models**\n\n```bash\n# Test code synthesis across 6 languages\naccelerate launch main.py \\\n  --model codellama/CodeLlama-7b-Instruct-hf \\\n  --tasks humanevalsynthesize-python,humanevalsynthesize-js \\\n  --prompt instruct \\\n  --max_length_generation 512 \\\n  --allow_code_execution\n```\n\n### Workflow 4: Compare Multiple Models\n\nBenchmark suite for model comparison.\n\n**Step 1: Create evaluation script**\n\n```bash\n#!/bin/bash\n# eval_models.sh\n\nMODELS=(\n  \"bigcode/starcoder2-7b\"\n  \"codellama/CodeLlama-7b-hf\"\n  \"deepseek-ai/deepseek-coder-6.7b-base\"\n)\nTASKS=\"humaneval,mbpp\"\n\nfor model in \"${MODELS[@]}\"; do\n  model_name=$(echo $model | tr '/' '-')\n  echo \"Evaluating $model\"\n\n  accelerate launch main.py \\\n    --model $model \\\n    --tasks $TASKS \\\n    --temperature 0.2 \\\n    --n_samples 20 \\\n    --batch_size 20 \\\n    --allow_code_execution \\\n    --metric_output_path results/${model_name}.json\ndone\n```\n\n**Step 2: Generate comparison table**\n\n```python\nimport json\nimport pandas as pd\n\nmodels = [\"bigcode-starcoder2-7b\", \"codellama-CodeLlama-7b-hf\", \"deepseek-ai-deepseek-coder-6.7b-base\"]\nresults = []\n\nfor model in models:\n    with open(f\"results/{model}.json\") as f:\n        data = json.load(f)\n        results.append({\n            \"Model\": model,\n            \"HumanEval pass@1\": f\"{data['humaneval']['pass@1']:.3f}\",\n            \"MBPP pass@1\": f\"{data['mbpp']['pass@1']:.3f}\"\n        })\n\ndf = pd.DataFrame(results)\nprint(df.to_markdown(index=False))\n```\n\n## When to Use vs Alternatives\n\n**Use BigCode Evaluation Harness when:**\n- Evaluating **code generation** models specifically\n- Need **multi-language** evaluation (18 languages via MultiPL-E)\n- Testing **functional correctness** with unit tests (pass@k)\n- Benchmarking for **BigCode/HuggingFace leaderboards**\n- Evaluating **fill-in-the-middle** (FIM) capabilities\n\n**Use alternatives instead:**\n- **lm-evaluation-harness**: General LLM benchmarks (MMLU, GSM8K, HellaSwag)\n- **EvalPlus**: Stricter HumanEval+/MBPP+ with more test cases\n- **SWE-bench**: Real-world GitHub issue resolution\n- **LiveCodeBench**: Contamination-free, continuously updated problems\n- **CodeXGLUE**: Code understanding tasks (clone detection, defect prediction)\n\n## Supported Benchmarks\n\n| Benchmark | Problems | Languages | Metric | Use Case |\n|-----------|----------|-----------|--------|----------|\n| HumanEval | 164 | Python | pass@k | Standard code completion |\n| HumanEval+ | 164 | Python | pass@k | Stricter evaluation (80× tests) |\n| MBPP | 500 | Python | pass@k | Entry-level problems |\n| MBPP+ | 399 | Python | pass@k | Stricter evaluation (35× tests) |\n| MultiPL-E | 164×18 | 18 languages | pass@k | Multi-language evaluation |\n| APPS | 10,000 | Python | pass@k | Competition-level |\n| DS-1000 | 1,000 | Python | pass@k | Data science (pandas, numpy, etc.) |\n| HumanEvalPack | 164×3×6 | 6 languages | pass@k | Synthesis/fix/explain |\n| Mercury | 1,889 | Python | Efficiency | Computational efficiency |\n\n## Common Issues\n\n**Issue: Different results than reported in papers**\n\nCheck these factors:\n```bash\n# 1. Verify n_samples (need 200 for accurate pass@k)\n--n_samples 200\n\n# 2. Check temperature (0.2 for greedy-ish, 0.8 for sampling)\n--temperature 0.8\n\n# 3. Verify task name matches exactly\n--tasks humaneval  # Not \"human_eval\" or \"HumanEval\"\n\n# 4. Check max_length_generation\n--max_length_generation 512  # Increase for longer problems\n```\n\n**Issue: CUDA out of memory**\n\n```bash\n# Use quantization\n--load_in_8bit\n# OR\n--load_in_4bit\n\n# Reduce batch size\n--batch_size 1\n\n# Set memory limit\n--max_memory_per_gpu \"20GiB\"\n```\n\n**Issue: Code execution hangs or times out**\n\nUse Docker for safe execution:\n```bash\n# Generate on host (no execution)\n--generation_only --save_generations\n\n# Evaluate in Docker\ndocker run ... --allow_code_execution --load_generations_path ...\n```\n\n**Issue: Low scores on instruction models**\n\nEnsure proper instruction formatting:\n```bash\n# Use instruction-specific tasks\n--tasks instruct-humaneval\n\n# Set instruction tokens for your model\n--instruction_tokens \"<s>[INST],</s>,[/INST]\"\n```\n\n**Issue: MultiPL-E language failures**\n\nUse the dedicated Docker image:\n```bash\ndocker pull ghcr.io/bigcode-project/evaluation-harness-multiple\n```\n\n## Command Reference\n\n| Argument | Default | Description |\n|----------|---------|-------------|\n| `--model` | - | HuggingFace model ID or local path |\n| `--tasks` | - | Comma-separated task names |\n| `--n_samples` | 1 | Samples per problem (200 for pass@k) |\n| `--temperature` | 0.2 | Sampling temperature |\n| `--max_length_generation` | 512 | Max tokens (prompt + generation) |\n| `--batch_size` | 1 | Batch size per GPU |\n| `--allow_code_execution` | False | Enable code execution (required) |\n| `--generation_only` | False | Generate without evaluation |\n| `--load_generations_path` | - | Load pre-generated solutions |\n| `--save_generations` | False | Save generated code |\n| `--metric_output_path` | results.json | Output file for metrics |\n| `--load_in_8bit` | False | 8-bit quantization |\n| `--load_in_4bit` | False | 4-bit quantization |\n| `--trust_remote_code` | False | Allow custom model code |\n| `--precision` | fp32 | Model precision (fp32/fp16/bf16) |\n\n## Hardware Requirements\n\n| Model Size | VRAM (fp16) | VRAM (4-bit) | Time (HumanEval, n=200) |\n|------------|-------------|--------------|-------------------------|\n| 7B | 14GB | 6GB | ~30 min (A100) |\n| 13B | 26GB | 10GB | ~1 hour (A100) |\n| 34B | 68GB | 20GB | ~2 hours (A100) |\n\n## Resources\n\n- **GitHub**: https://github.com/bigcode-project/bigcode-evaluation-harness\n- **Documentation**: https://github.com/bigcode-project/bigcode-evaluation-harness/tree/main/docs\n- **BigCode Leaderboard**: https://huggingface.co/spaces/bigcode/bigcode-models-leaderboard\n- **HumanEval Dataset**: https://huggingface.co/datasets/openai/openai_humaneval\n- **MultiPL-E**: https://github.com/nuprl/MultiPL-E","author":"@Orchestra-Research","ownerProfile":null,"authorContacts":null,"sourceUrl":"https://github.com/Orchestra-Research/AI-Research-SKILLs/tree/main/11-evaluation/bigcode-evaluation-harness","license":"MIT","category":"coding","lang":"en","tokens":3087,"stars":0,"calls30d":2,"claimed":false,"visibility":"public","origin":"crawler","version":"0.1.0","createdAt":"2026-08-22","updatedAt":"2026-08-22","files":[{"path":"references/benchmarks.md","size":9168,"sha256":"50c5640497be1861dc69ef886f2e4f4684162c28d57ee928111f1c4c71e4be2d"},{"path":"references/custom-tasks.md","size":10378,"sha256":"4f868c5c84d1889fbd6e3f09297ce439be70a9af6f8694bd274481ab09677272"},{"path":"references/issues.md","size":8085,"sha256":"09832ab736fec29fc949101f81b806c0c6fc366124eada17967aac72ea5a63a1"}],"requires":{"mcp":[],"tools":[]},"safety":{"flags":[],"scannedAt":"2026-08-22","hasScripts":false,"networkEndpoints":["download.pytorch.org","huggingface.co"]}}