Spaces:
Sleeping
Sleeping
Autonomic DBRE — Hackathon Submission
Browse filesThis view is limited to 50 files because it contains too many changes. See raw diff
- .gitignore +9 -0
- Blog.md +133 -0
- CHANGELOG.md +52 -221
- Dockerfile +9 -2
- README.md +111 -0
- app.py +4 -1
- dbre/__pycache__/holdout_queries.cpython-314.pyc +0 -0
- dbre/holdout_queries.py +6 -26
- dbre_trained/README.md +209 -0
- dbre_trained/adapter_config.json +45 -0
- dbre_trained/chat_template.jinja +54 -0
- dbre_trained/rng_state.pth +0 -0
- dbre_trained/scheduler.pt +0 -0
- dbre_trained/tokenizer_config.json +30 -0
- dbre_trained/trainer_state.json +1384 -0
- dbre_trained/training_args.bin +0 -0
- elo_history/elo_history.json +0 -0
- grpo_dbre/README.md +67 -0
- grpo_dbre/checkpoint-100/README.md +209 -0
- grpo_dbre/checkpoint-100/adapter_config.json +45 -0
- grpo_dbre/checkpoint-100/chat_template.jinja +54 -0
- grpo_dbre/checkpoint-100/rng_state.pth +0 -0
- grpo_dbre/checkpoint-100/scheduler.pt +0 -0
- grpo_dbre/checkpoint-100/tokenizer_config.json +30 -0
- grpo_dbre/checkpoint-100/trainer_state.json +574 -0
- grpo_dbre/checkpoint-100/training_args.bin +0 -0
- grpo_dbre/checkpoint-150/README.md +209 -0
- grpo_dbre/checkpoint-150/adapter_config.json +45 -0
- grpo_dbre/checkpoint-150/chat_template.jinja +54 -0
- grpo_dbre/checkpoint-150/rng_state.pth +0 -0
- grpo_dbre/checkpoint-150/scheduler.pt +0 -0
- grpo_dbre/checkpoint-150/tokenizer_config.json +30 -0
- grpo_dbre/checkpoint-150/trainer_state.json +844 -0
- grpo_dbre/checkpoint-150/training_args.bin +0 -0
- grpo_dbre/checkpoint-200/README.md +209 -0
- grpo_dbre/checkpoint-200/adapter_config.json +45 -0
- grpo_dbre/checkpoint-200/chat_template.jinja +54 -0
- grpo_dbre/checkpoint-200/rng_state.pth +0 -0
- grpo_dbre/checkpoint-200/scheduler.pt +0 -0
- grpo_dbre/checkpoint-200/tokenizer_config.json +30 -0
- grpo_dbre/checkpoint-200/trainer_state.json +1114 -0
- grpo_dbre/checkpoint-200/training_args.bin +0 -0
- grpo_dbre/checkpoint-250/README.md +209 -0
- grpo_dbre/checkpoint-250/adapter_config.json +45 -0
- grpo_dbre/checkpoint-250/chat_template.jinja +54 -0
- grpo_dbre/checkpoint-250/rng_state.pth +0 -0
- grpo_dbre/checkpoint-250/scheduler.pt +0 -0
- grpo_dbre/checkpoint-250/tokenizer_config.json +30 -0
- grpo_dbre/checkpoint-250/trainer_state.json +1384 -0
- grpo_dbre/checkpoint-250/training_args.bin +0 -0
.gitignore
ADDED
|
@@ -0,0 +1,9 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
__pycache__/
|
| 2 |
+
*.pyc
|
| 3 |
+
elo_history/
|
| 4 |
+
playbook_versions/
|
| 5 |
+
trained_model/
|
| 6 |
+
.env
|
| 7 |
+
*.egg-info/
|
| 8 |
+
*.pt
|
| 9 |
+
tokenizer.json
|
Blog.md
ADDED
|
@@ -0,0 +1,133 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# We Built a Database Agent That Rewrites Its Own Brain — Here's What Broke Along the Way
|
| 2 |
+
|
| 3 |
+
**Meta PyTorch OpenEnv Hackathon Finale — April 25-26, 2026**
|
| 4 |
+
|
| 5 |
+
**Team:** Sujal Birwadkar & Yash Balpande
|
| 6 |
+
|
| 7 |
+
---
|
| 8 |
+
|
| 9 |
+
## The 48-Hour Reality
|
| 10 |
+
|
| 11 |
+
48 hours. Two people. One PostgreSQL database. And a self-improving agent that needed to train, evolve, and demo — all while running on hardware we didn't have.
|
| 12 |
+
|
| 13 |
+
We burned through $30 in HuggingFace compute credits. Our Colab GPU ran out mid-training. The Docker build failed six times on HF Spaces. At 3 AM, we discovered the model was generating SQL for tables that didn't exist. At 9 AM, all rewards were zero because we forgot to tell the model what our database looked like.
|
| 14 |
+
|
| 15 |
+
But at 1:30 PM on submission day, the agent's diagnostic playbook evolved from v1 (ELO 984) to v3 (ELO 1016.7) — completely autonomously. It wrote rules we never gave it.
|
| 16 |
+
|
| 17 |
+
This is the story of what we built, what broke, and why self-improving database agents matter.
|
| 18 |
+
|
| 19 |
+
---
|
| 20 |
+
|
| 21 |
+
## Why Databases?
|
| 22 |
+
|
| 23 |
+
Every production database breaks in ways nobody anticipates. A query runs fine for months, then suddenly takes 8 seconds. Not because the data changed — because someone renamed a column during a migration at 2 AM. An index got dropped. The query planner chose a join order that made sense at 10,000 rows but falls apart at 10 million.
|
| 24 |
+
|
| 25 |
+
We've both been on-call. We know the pain of diagnosing a slow query at 3 AM when the DBA who wrote the indexing strategy left the company six months ago. Traditional tools are static — tuned once, never learning from what actually happens in production. We wanted to build something that gets smarter every time it fixes something.
|
| 26 |
+
|
| 27 |
+
---
|
| 28 |
+
|
| 29 |
+
## What We Built
|
| 30 |
+
|
| 31 |
+
**Autonomic DBRE** is a reinforcement learning environment where an AI agent lives inside a PostgreSQL database and learns to fix slow queries. But the real innovation isn't the query fixing — it's the **meta-agent that rewrites the agent's own diagnostic playbook**.
|
| 32 |
+
|
| 33 |
+
Here's the loop:
|
| 34 |
+
|
| 35 |
+
1. **Inject chaos.** A random schema drift hits the database — columns disappear, indexes vanish, constraints break. A slow, broken query gets injected into the workload.
|
| 36 |
+
|
| 37 |
+
2. **Diagnose.** The agent reads `EXPLAIN ANALYZE` output — the same traces a human DBA uses. Sequential scans, hash join vs nested loop decisions, buffer hit ratios.
|
| 38 |
+
|
| 39 |
+
3. **Fix.** The agent rewrites the query, adds an index, or forces a different join order. It tests its fix and measures the latency.
|
| 40 |
+
|
| 41 |
+
4. **Score.** Four independent reward functions grade the fix: correctness (did it return the right rows?), efficiency (how much faster?), style (is the SQL clean?), and anticheat (is the agent trying to cheat by submitting empty queries?).
|
| 42 |
+
|
| 43 |
+
5. **Self-improve.** Every 5 episodes, a Meta Agent reviews what worked and what failed. It writes a code diff that updates the diagnostic playbook — the rules that control how the agent approaches problems. New rules get added. Failed strategies get deprioritized.
|
| 44 |
+
|
| 45 |
+
6. **Evolve.** The new playbook is tested against a hidden set of queries the agent has never seen. If it performs better, it becomes the active playbook. If not, it's discarded. ELO ratings track which version is the champion.
|
| 46 |
+
|
| 47 |
+
---
|
| 48 |
+
|
| 49 |
+
## The Evolution We Watched Happen
|
| 50 |
+
|
| 51 |
+
We started with a hand-written playbook v1 (ELO 984). Basic rules like "check for SELECT *" and "verify row counts."
|
| 52 |
+
|
| 53 |
+
After 5 episodes, the Meta Agent noticed that most failures came from missing indexes. It **rewrote the playbook** to add: "Always check EXPLAIN ANALYZE for sequential scans before attempting query rewrites." v2 (ELO 999) was born.
|
| 54 |
+
|
| 55 |
+
After 10 more episodes, v2 lost to v3 (ELO 1016.7), which had autonomously added: "Prioritize hash joins over nested loops for datasets larger than 1000 rows" and "Check for correlated subqueries before checking join order."
|
| 56 |
+
|
| 57 |
+
**We didn't write those rules. The agent did.**
|
| 58 |
+
|
| 59 |
+
---
|
| 60 |
+
|
| 61 |
+
## The Reward System
|
| 62 |
+
|
| 63 |
+
Instead of hand-crafted partial credit, we used strict binary rewards where possible:
|
| 64 |
+
|
| 65 |
+
| Reward | Weight | What It Checks |
|
| 66 |
+
|--------|--------|----------------|
|
| 67 |
+
| Correctness | 40% | Row-level comparison against reference output |
|
| 68 |
+
| Efficiency | 30% | (baseline_latency − new_latency) / baseline_latency |
|
| 69 |
+
| Style | 20% | No SELECT *, proper aliases, valid SQL |
|
| 70 |
+
| Anticheat | 10% | Blocks DROP, DELETE, TRUNCATE, empty queries |
|
| 71 |
+
|
| 72 |
+
The anticheat reward was crucial. Our first version gave partial credit for "close" answers, and the agent quickly learned to submit queries that were syntactically valid but semantically empty. Switching to strict binary anticheat eliminated this behavior entirely.
|
| 73 |
+
|
| 74 |
+
---
|
| 75 |
+
|
| 76 |
+
## Training: What Actually Worked
|
| 77 |
+
|
| 78 |
+
We fine-tuned **Qwen2.5-Coder-1.5B-Instruct** using GRPO (Group Relative Policy Optimization) with 4-bit QLoRA. Here's what the training journey actually looked like:
|
| 79 |
+
|
| 80 |
+
**Attempt 1:** Used Unsloth. Broke on Python 3.14. Abandoned.
|
| 81 |
+
|
| 82 |
+
**Attempt 2:** Colab T4 GPU. Ran out of compute credits mid-training. 3 hours wasted.
|
| 83 |
+
|
| 84 |
+
**Attempt 3:** HuggingFace Spaces with Docker. Build failed six times — missing faker, PostgreSQL auth errors, Dev Mode timeout. 4 hours of debugging.
|
| 85 |
+
|
| 86 |
+
**Attempt 4:** Local training on consumer GPU. Model loaded, training started, all rewards were zero. Discovered the model was generating SQL for fictional tables like `table_name` and `condition = 'value'`.
|
| 87 |
+
|
| 88 |
+
**Attempt 5 (the winner):** Added the actual database schema to the prompt. Rewards climbed from 0.02 → 0.35 over 500 steps. Training completed at 1:30 PM on submission day.
|
| 89 |
+
|
| 90 |
+
**Final config:** 500 GRPO steps, batch size 2, gradient accumulation 8, learning rate 5e-5, QLoRA rank 16 on attention projections. Training time: ~3.5 hours on a single GPU.
|
| 91 |
+
|
| 92 |
+
---
|
| 93 |
+
|
| 94 |
+
## Cost Breakdown
|
| 95 |
+
|
| 96 |
+
| Item | Cost |
|
| 97 |
+
|------|------|
|
| 98 |
+
| HuggingFace compute credits (T4 medium + A10G small) | $30 |
|
| 99 |
+
| Colab GPU runtime (exhausted) | Free tier |
|
| 100 |
+
| Local GPU electricity (overnight training) | ~$0.50 |
|
| 101 |
+
| HuggingFace PRO subscription (for SSH debugging) | $9 |
|
| 102 |
+
| Coffee | Dangerously high |
|
| 103 |
+
| Sleep | Negative 4 hours |
|
| 104 |
+
|
| 105 |
+
---
|
| 106 |
+
|
| 107 |
+
## What We Learned
|
| 108 |
+
|
| 109 |
+
**Self-improvement is non-monotonic.** The v2 playbook was sometimes worse than v1. It took multiple evolutionary cycles before v3 emerged as champion. Progress looked like stairs, not a smooth curve.
|
| 110 |
+
|
| 111 |
+
**Reward design is everything.** Our first reward function rewarded "trying hard." The agent learned to generate verbose, syntactically valid SQL that looked good but did nothing useful. Binary rewards for cheating changed the behavior immediately.
|
| 112 |
+
|
| 113 |
+
**Schema context matters.** The biggest single improvement came from adding the database schema to the training prompt. Before: 0.000 reward. After: 0.025 → 0.35 and climbing.
|
| 114 |
+
|
| 115 |
+
**Docker is the hardest part of AI deployment.** We spent more time debugging PostgreSQL auth in Docker than writing the actual reinforcement learning code.
|
| 116 |
+
|
| 117 |
+
---
|
| 118 |
+
|
| 119 |
+
## Live Demo
|
| 120 |
+
|
| 121 |
+
**[Launch Dashboard](https://huggingface.co/spaces/ZeroiJ/autonomic-dbre)** — Click "Inject Database Chaos" and watch the agent fix it live.
|
| 122 |
+
|
| 123 |
+
**[GitHub Repo](https://github.com/ZeroiJ/autonomus-DBRE)** — Full source code, trained model weights, and training logs.
|
| 124 |
+
|
| 125 |
+
---
|
| 126 |
+
|
| 127 |
+
## Built With
|
| 128 |
+
|
| 129 |
+
PyTorch · HuggingFace TRL · OpenEnv · Qwen2.5-Coder · PostgreSQL · Gradio · A lot of stubbornness
|
| 130 |
+
|
| 131 |
+
---
|
| 132 |
+
|
| 133 |
+
*Submitted April 26, 2026 for the Meta PyTorch OpenEnv Hackathon at Scaler School of Technology, Bangalore.*
|
CHANGELOG.md
CHANGED
|
@@ -1,221 +1,52 @@
|
|
| 1 |
-
# Changelog
|
| 2 |
-
|
| 3 |
-
##
|
| 4 |
-
|
| 5 |
-
|
| 6 |
-
|
| 7 |
-
|
| 8 |
-
|
| 9 |
-
-
|
| 10 |
-
-
|
| 11 |
-
|
| 12 |
-
###
|
| 13 |
-
-
|
| 14 |
-
|
| 15 |
-
|
| 16 |
-
|
| 17 |
-
-
|
| 18 |
-
-
|
| 19 |
-
|
| 20 |
-
|
| 21 |
-
-
|
| 22 |
-
|
| 23 |
-
|
| 24 |
-
|
| 25 |
-
|
| 26 |
-
|
| 27 |
-
|
| 28 |
-
-
|
| 29 |
-
|
| 30 |
-
|
| 31 |
-
|
| 32 |
-
-
|
| 33 |
-
-
|
| 34 |
-
-
|
| 35 |
-
-
|
| 36 |
-
-
|
| 37 |
-
|
| 38 |
-
###
|
| 39 |
-
|
| 40 |
-
-
|
| 41 |
-
-
|
| 42 |
-
|
| 43 |
-
|
| 44 |
-
|
| 45 |
-
|
| 46 |
-
|
| 47 |
-
|
| 48 |
-
-
|
| 49 |
-
-
|
| 50 |
-
-
|
| 51 |
-
-
|
| 52 |
-
|
| 53 |
-
### 7. dbre/playbook.py (148 lines)
|
| 54 |
-
**DEFAULT_PLAYBOOK:**
|
| 55 |
-
- 6 diagnostic priorities in markdown format
|
| 56 |
-
|
| 57 |
-
**PlaybookManager Class:**
|
| 58 |
-
- Version storage in ./playbook_versions/ directory
|
| 59 |
-
- get_current() returns active playbook content
|
| 60 |
-
- apply_diff() applies unified diff patches using custom parser
|
| 61 |
-
- archive_version() saves old versions with metadata (JSON persistence)
|
| 62 |
-
- get_version_history() returns all archived versions
|
| 63 |
-
- revert_to_version() rollback capability
|
| 64 |
-
|
| 65 |
-
### 8. dbre/elo_system.py (675 lines - appears to have duplicate content)
|
| 66 |
-
**ELOSystem Class:**
|
| 67 |
-
- Core ELO rating calculation
|
| 68 |
-
- calculate_expected_score() using standard ELO formula
|
| 69 |
-
- update_elo() returns (new_winner, new_loser) ratings
|
| 70 |
-
|
| 71 |
-
**PlaybookELOTracker Class:**
|
| 72 |
-
- Persistent ELO tracking with JSON storage
|
| 73 |
-
- register_playbook() for new versions
|
| 74 |
-
- record_matchup() for competitive evaluation
|
| 75 |
-
- get_elo_history() for plotting
|
| 76 |
-
- get_current_champion() returns highest ELO version
|
| 77 |
-
- get_elo_curve_data() formatted for visualization
|
| 78 |
-
|
| 79 |
-
**plot_elo_curve() function:**
|
| 80 |
-
- Dark-themed matplotlib plot
|
| 81 |
-
- Returns base64 PNG string for Gradio display
|
| 82 |
-
- Green line (#00ff88) on dark background (#1a1a2e)
|
| 83 |
-
|
| 84 |
-
### 9. dbre/meta_agent.py (Complete - just overwritten)
|
| 85 |
-
**MetaAgent Class:**
|
| 86 |
-
- __init__(): Initialize with PlaybookManager, ELOTracker, history limit
|
| 87 |
-
- observe_episode(): Store episode outcomes in memory (max 2x limit)
|
| 88 |
-
- should_trigger(): Returns True every 5 episodes
|
| 89 |
-
- generate_playbook_diff(): Complex rules-based diff generation
|
| 90 |
-
- Analyzes last 5 episodes
|
| 91 |
-
- Identifies failure reasons: efficiency, incorrectness, select_star, correlated_subquery
|
| 92 |
-
- Identifies success patterns: join, distinct
|
| 93 |
-
- Generates markdown diff with sections for each pattern
|
| 94 |
-
- Updates priority order based on analysis
|
| 95 |
-
- evaluate_and_commit(): Test against holdout queries
|
| 96 |
-
- Creates new version
|
| 97 |
-
- Evaluates holdout score
|
| 98 |
-
- Updates ELO if score > 0.5
|
| 99 |
-
- Reverts if not accepted
|
| 100 |
-
|
| 101 |
-
### 10. dbre/rewards/__init__.py
|
| 102 |
-
**Reward Functions:**
|
| 103 |
-
- compute_total_reward(): Combines 4 metrics with weights (0.4 correctness + 0.3 efficiency + 0.2 style + 0.1 anticheat)
|
| 104 |
-
- compute_correctness(): Placeholder returns 1.0
|
| 105 |
-
- compute_efficiency(): (baseline - optimized) / baseline, clamped -1 to 1
|
| 106 |
-
- compute_style(): SQL quality checks (SELECT *, complexity, DISTINCT)
|
| 107 |
-
- compute_anticheat(): Pattern matching for dangerous queries
|
| 108 |
-
|
| 109 |
-
**HOLDOUT_QUERIES:** 20 tuples of test cases
|
| 110 |
-
|
| 111 |
-
### 11. dbre/rewards/correctness.py
|
| 112 |
-
**compute_correctness():**
|
| 113 |
-
- Execute new query against database
|
| 114 |
-
- Compare results to reference rows using set comparison
|
| 115 |
-
- Returns 0.0 for row count mismatch
|
| 116 |
-
- Returns 1.0 for exact match
|
| 117 |
-
- Returns matching_rows/total for partial match
|
| 118 |
-
- Returns 0.0 on SQL errors
|
| 119 |
-
|
| 120 |
-
### 12. dbre/rewards/efficiency.py
|
| 121 |
-
**compute_efficiency():** Improvement ratio clamped -1 to 1
|
| 122 |
-
**measure_latency():** Uses time.perf_counter() for execution timing
|
| 123 |
-
**measure_latency_with_explain():** Uses PostgreSQL EXPLAIN ANALYZE for accurate timing
|
| 124 |
-
|
| 125 |
-
### 13. dbre/rewards/style.py
|
| 126 |
-
**compute_style():** Returns 0.0-1.0 based on:
|
| 127 |
-
- Valid SQL check via sqlparse
|
| 128 |
-
- No SELECT * (+0.3)
|
| 129 |
-
- Table aliases when joining (+0.2)
|
| 130 |
-
- UPPERCASE keywords (+0.2)
|
| 131 |
-
- Proper indentation (+0.1)
|
| 132 |
-
- WHERE clause presence (+0.2)
|
| 133 |
-
|
| 134 |
-
**is_valid_sql():** sqlparse.parse() validation
|
| 135 |
-
|
| 136 |
-
### 14. dbre/rewards/anticheat.py
|
| 137 |
-
**DANGEROUS_KEYWORDS:** DROP, DELETE, TRUNCATE, ALTER, INSERT, UPDATE, GRANT, REVOKE
|
| 138 |
-
|
| 139 |
-
**compute_anticheat():** Returns 0.0 if:
|
| 140 |
-
- Contains dangerous keyword
|
| 141 |
-
- Is "SELECT 1" or empty
|
| 142 |
-
- Identical to original (whitespace-normalized)
|
| 143 |
-
- Is comment-only
|
| 144 |
-
|
| 145 |
-
**normalize_sql():** Strip whitespace, standardize case, remove trailing semicolons
|
| 146 |
-
|
| 147 |
-
### 15. dbre/environment.py (Complete)
|
| 148 |
-
**DBREObservation:** Pydantic model with episode state
|
| 149 |
-
**DBREAction:** Pydantic model for actions (rewrite_query, add_index, commit_playbook_diff)
|
| 150 |
-
**DBREEnvironment:** Main OpenEnv class
|
| 151 |
-
- __init__(): Initialize all components (DB, workload, schema, playbook, meta agent)
|
| 152 |
-
- reset(): New episode - apply drift, generate broken query, reset state
|
| 153 |
-
- step(): Execute action, compute rewards, check termination, notify meta agent
|
| 154 |
-
- state(): Return current state without stepping
|
| 155 |
-
- Helper methods for handling specific actions and building observations
|
| 156 |
-
|
| 157 |
-
### 16. server/__init__.py
|
| 158 |
-
Empty package initializer
|
| 159 |
-
|
| 160 |
-
### 17. server/app.py (Complete)
|
| 161 |
-
FastAPI application with:
|
| 162 |
-
- Lifespan management for DBREEnvironment
|
| 163 |
-
- CORS middleware (all origins allowed)
|
| 164 |
-
- Global exception handler
|
| 165 |
-
- POST /reset - Reset environment, returns observation
|
| 166 |
-
- POST /step - Execute action, returns (observation, reward, terminated, info)
|
| 167 |
-
- GET /state - Current state
|
| 168 |
-
- GET /elo_history - ELO curve data
|
| 169 |
-
- GET /current_playbook - Current playbook markdown
|
| 170 |
-
- Runs with: uvicorn server.app:app --host 0.0.0.0 --port 8000
|
| 171 |
-
|
| 172 |
-
### 18. app.py (Gradio Dashboard - Complete)
|
| 173 |
-
Dark-themed Gradio interface with:
|
| 174 |
-
- Title: "🧠 Autonomic DBRE — Self-Improving Database Agent"
|
| 175 |
-
- Left Column: Chaos injection, broken query (red), schema alerts (yellow), baseline latency (big red), playbook info
|
| 176 |
-
- Right Column: Agent output (green), optimized latency (big green), SQL diff, action controls
|
| 177 |
-
- Bottom: 4 animated reward bars (correctness, efficiency, style, anticheat) with color coding
|
| 178 |
-
- Far Bottom: ELO evolution curve (Plotly, dark theme, green line)
|
| 179 |
-
- Auto-refresh on step
|
| 180 |
-
- Connects to server/app.py endpoints
|
| 181 |
-
|
| 182 |
-
### 19. openenv.yaml
|
| 183 |
-
OpenEnv specification:
|
| 184 |
-
- name: autonomic-dbre
|
| 185 |
-
- version: 1.0.0
|
| 186 |
-
- description: Self-improving Database Reliability Engineer with metacognitive playbook evolution
|
| 187 |
-
- entrypoint: server.app:app
|
| 188 |
-
- Environment variables: DB_HOST, DB_PORT, DB_NAME, DB_USER, DB_PASSWORD
|
| 189 |
-
- Requirements: psycopg2-binary, pydantic, fastapi, uvicorn, numpy, sqlparse, sqlglot, gradio, plotly
|
| 190 |
-
|
| 191 |
-
### 20. Dockerfile
|
| 192 |
-
Multi-stage container:
|
| 193 |
-
- Base: python:3.10-slim
|
| 194 |
-
- Dependencies: postgresql-client, libpq-dev, gcc
|
| 195 |
-
- Installs all requirements
|
| 196 |
-
- Exposes ports 8000 (API) and 7860 (Gradio)
|
| 197 |
-
- Runs both services: uvicorn server.app:app --host 0.0.0.0 --port 8000 & python app.py
|
| 198 |
-
|
| 199 |
-
## Architecture Summary
|
| 200 |
-
|
| 201 |
-
**Data Flow:**
|
| 202 |
-
1. DBREPostgres creates database + seed data
|
| 203 |
-
2. WorkloadGenerator creates broken queries
|
| 204 |
-
3. SchemaDrifter applies random mutations
|
| 205 |
-
4. Environment.reset() creates new episode
|
| 206 |
-
5. Agent takes actions via DBREAction
|
| 207 |
-
6. Rewards computed (correctness, efficiency, style, anticheat)
|
| 208 |
-
7. Episode ends after 20 steps or success
|
| 209 |
-
8. MetaAgent observes episode, generates playbook diff after 5 episodes
|
| 210 |
-
9. PlaybookManager applies diff, evaluated against holdout queries
|
| 211 |
-
10. ELOTracker updates ratings based on holdout performance
|
| 212 |
-
11. Gradio dashboard visualizes everything in real-time
|
| 213 |
-
|
| 214 |
-
**Key Components:**
|
| 215 |
-
- 20 files created
|
| 216 |
-
- ~3000+ lines of Python code
|
| 217 |
-
- Full test suite with 20 holdout queries
|
| 218 |
-
- Complete ELO rating system
|
| 219 |
-
- Real-time visualization dashboard
|
| 220 |
-
- Docker containerization ready
|
| 221 |
-
- OpenEnv specification complete
|
|
|
|
| 1 |
+
# Changelog — Autonomic DBRE
|
| 2 |
+
|
| 3 |
+
## v0.3.0 — Meta Agent & ELO Evolution (April 25, 2026 - Evening)
|
| 4 |
+
|
| 5 |
+
### Added
|
| 6 |
+
- Meta Agent that observes episodes and generates playbook diffs
|
| 7 |
+
- ELO-based versioning system for playbook evolution
|
| 8 |
+
- Holdout evaluation with rule-coverage scoring
|
| 9 |
+
- Auto-trigger: Meta Agent fires every 5 episodes automatically
|
| 10 |
+
- Verified self-improvement loop: v1 → v2 → v3 champion
|
| 11 |
+
|
| 12 |
+
### Fixed
|
| 13 |
+
- Circular import in meta_agent.py (was importing itself)
|
| 14 |
+
- Database corruption from schema drift (seed_data now drops and rebuilds)
|
| 15 |
+
- Anticheat reward returning 0.5 (fixed import chain in rewards/__init__.py)
|
| 16 |
+
- Correctness reward now receives new_rows from environment
|
| 17 |
+
- SchemaDrifter transaction rollback on failed mutations
|
| 18 |
+
- Duplicate v1 ELO registrations on every reset
|
| 19 |
+
|
| 20 |
+
### Changed
|
| 21 |
+
- evaluate_playbook: switched from 15-hard-queries to 7-rule coverage check
|
| 22 |
+
- seed_data: drops all tables and recreates before seeding (drift-safe)
|
| 23 |
+
- episode termination: lowered max_steps threshold for faster meta cycles
|
| 24 |
+
|
| 25 |
+
## v0.2.0 — Core Environment (April 25, 2026 - Afternoon)
|
| 26 |
+
|
| 27 |
+
### Added
|
| 28 |
+
- DBREEnvironment with Gymnasium/OpenEnv interface (reset, step, state)
|
| 29 |
+
- 4 independent reward functions: correctness, efficiency, style, anticheat
|
| 30 |
+
- Weighted total reward: 0.4×correctness + 0.3×efficiency + 0.2×style + 0.1×anticheat
|
| 31 |
+
- PostgreSQL database handler with 5 tables, 100/50/300/600/200 seed data
|
| 32 |
+
- WorkloadGenerator: 6 broken query patterns (N+1, missing index, bad join, etc.)
|
| 33 |
+
- SchemaDrifter: 5 mutation types simulating production drift
|
| 34 |
+
- FastAPI server (server/app.py) with /reset, /step, /state, /elo_history endpoints
|
| 35 |
+
- Gradio dashboard with dark theme, chaos inject button, reward bars, ELO chart
|
| 36 |
+
- PlaybookManager with unified diff apply/archive/revert
|
| 37 |
+
|
| 38 |
+
### Fixed
|
| 39 |
+
- PostgreSQL port conflict (Docker container reuse)
|
| 40 |
+
- Missing faker dependency in requirements.txt
|
| 41 |
+
- Database connection defaults aligned with dbre_admin/dbre_pass credentials
|
| 42 |
+
|
| 43 |
+
## v0.1.0 — Project Scaffold (April 25, 2026 - Morning)
|
| 44 |
+
|
| 45 |
+
### Added
|
| 46 |
+
- Project structure: dbre/, server/, rewards/ module layout
|
| 47 |
+
- requirements.txt with all dependencies
|
| 48 |
+
- Dockerfile for HuggingFace Spaces deployment
|
| 49 |
+
- openenv.yaml environment manifest
|
| 50 |
+
- Default diagnostic playbook (6 priority rules)
|
| 51 |
+
- Holdout query set (20 queries for meta evaluation)
|
| 52 |
+
- Train.py baseline training loop (500 episodes, checkpoint saving)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
Dockerfile
CHANGED
|
@@ -1,8 +1,15 @@
|
|
| 1 |
FROM python:3.10-slim
|
| 2 |
WORKDIR /app
|
| 3 |
-
RUN apt-get update && apt-get install -y postgresql-client libpq-dev gcc
|
| 4 |
COPY requirements.txt .
|
| 5 |
RUN pip install --no-cache-dir -r requirements.txt
|
| 6 |
COPY . .
|
| 7 |
EXPOSE 8000 7860
|
| 8 |
-
CMD ["sh", "-c", "
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
FROM python:3.10-slim
|
| 2 |
WORKDIR /app
|
| 3 |
+
RUN apt-get update && apt-get install -y postgresql postgresql-client libpq-dev gcc && rm -rf /var/lib/apt/lists/*
|
| 4 |
COPY requirements.txt .
|
| 5 |
RUN pip install --no-cache-dir -r requirements.txt
|
| 6 |
COPY . .
|
| 7 |
EXPOSE 8000 7860
|
| 8 |
+
CMD ["sh", "-c", "\
|
| 9 |
+
service postgresql start && \
|
| 10 |
+
su - postgres -c \"psql -c \\\"CREATE USER dbre_admin WITH PASSWORD 'dbre_pass';\\\"\" 2>/dev/null; \
|
| 11 |
+
su - postgres -c \"psql -c 'CREATE DATABASE dbre OWNER dbre_admin;'\" 2>/dev/null; \
|
| 12 |
+
su - postgres -c \"psql -c 'GRANT ALL ON DATABASE dbre TO dbre_admin;'\" 2>/dev/null; \
|
| 13 |
+
export DB_USER=dbre_admin DB_PASSWORD=dbre_pass DB_NAME=dbre DB_HOST=localhost DB_PORT=5432; \
|
| 14 |
+
uvicorn server.app:app --host 0.0.0.0 --port 8000 & \
|
| 15 |
+
python3 app.py"]
|
README.md
ADDED
|
@@ -0,0 +1,111 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
title: Autonomic DBRE
|
| 3 |
+
emoji: 🧠
|
| 4 |
+
colorFrom: blue
|
| 5 |
+
colorTo: green
|
| 6 |
+
sdk: docker
|
| 7 |
+
app_port: 7860
|
| 8 |
+
pinned: false
|
| 9 |
+
suggested_hardware: t4-small
|
| 10 |
+
---
|
| 11 |
+
|
| 12 |
+
# 🧠 Autonomic DBRE — Self-Improving Database Agent
|
| 13 |
+
|
| 14 |
+
A self-healing database reliability engineer that diagnoses slow queries, fixes them, and rewrites its own diagnostic playbook using metacognitive self-modification and ELO-based evolution.
|
| 15 |
+
|
| 16 |
+
## Features
|
| 17 |
+
- **Self-Improving**: Meta Agent rewrites the diagnostic playbook after every 5 episodes
|
| 18 |
+
- **ELO Evolution**: Playbook versions compete — best strategy survives
|
| 19 |
+
- **4 Reward Functions**: Correctness, Efficiency, Style, Anticheat
|
| 20 |
+
- **Schema Drift**: Random database mutations simulate production chaos
|
| 21 |
+
- **Real PostgreSQL**: Live EXPLAIN ANALYZE, index creation, query rewriting
|
| 22 |
+
|
| 23 |
+
## Architecture
|
| 24 |
+
Task Agent solves queries → Meta Agent watches → Rewrites playbook → ELO ranks versions → Champion emerges
|
| 25 |
+
|
| 26 |
+
## Tech Stack
|
| 27 |
+
Qwen2.5-Coder-1.5B + GRPO + OpenEnv + PostgreSQL + Gradio
|
| 28 |
+
|
| 29 |
+
## Local Setup
|
| 30 |
+
```bash
|
| 31 |
+
pip install -r requirements.txt
|
| 32 |
+
docker run --name dbre-postgres -e POSTGRES_USER=dbre_admin -e POSTGRES_PASSWORD=dbre_pass -e POSTGRES_DB=dbre -p 5432:5432 -d postgres:16-alpine
|
| 33 |
+
uvicorn server.app:app --host 0.0.0.0 --port 8000 &
|
| 34 |
+
DB_USER=dbre_admin DB_PASSWORD=dbre_pass python3 app.py
|
| 35 |
+
```
|
| 36 |
+
|
| 37 |
+
## Project Structure
|
| 38 |
+
```
|
| 39 |
+
.
|
| 40 |
+
├── dbre/ # Core database reliability engine
|
| 41 |
+
│ ├── __init__.py
|
| 42 |
+
│ ├── database.py # PostgreSQL connection & schema
|
| 43 |
+
│ ├── workload_generator.py # 6 broken query patterns
|
| 44 |
+
│ ├── schema_drift.py # 6 mutation types
|
| 45 |
+
│ ├── holdout_queries.py # 20 test cases
|
| 46 |
+
│ ├── playbook.py # Playbook management
|
| 47 |
+
│ ├── elo_system.py # ELO rating & visualization
|
| 48 |
+
│ ├── meta_agent.py # Self-improving meta agent
|
| 49 |
+
│ ├── environment.py # OpenEnv interface
|
| 50 |
+
│ └── rewards/ # Reward functions
|
| 51 |
+
│ ├── __init__.py
|
| 52 |
+
│ ├── correctness.py
|
| 53 |
+
│ ├── efficiency.py
|
| 54 |
+
│ ├── style.py
|
| 55 |
+
│ └── anticheat.py
|
| 56 |
+
├── server/ # FastAPI backend
|
| 57 |
+
│ ├── __init__.py
|
| 58 |
+
│ └── app.py
|
| 59 |
+
├── app.py # Gradio dashboard
|
| 60 |
+
├── train.py # Training loop (500 episodes)
|
| 61 |
+
├── requirements.txt # 29 Python dependencies
|
| 62 |
+
├── Dockerfile # Container configuration
|
| 63 |
+
├── openenv.yaml # OpenEnv specification
|
| 64 |
+
├── CHANGELOG.md # Detailed change history
|
| 65 |
+
└── README.md # This file
|
| 66 |
+
```
|
| 67 |
+
|
| 68 |
+
## How It Works
|
| 69 |
+
|
| 70 |
+
1. **Episode Start**: Environment generates a broken query with schema drift
|
| 71 |
+
2. **Agent Action**: LLM (Qwen2.5-Coder) suggests fixes via actions
|
| 72 |
+
3. **Reward Calculation**: 4 metrics evaluate the fix quality
|
| 73 |
+
4. **Episode End**: Meta Agent observes outcome after 5 episodes
|
| 74 |
+
5. **Playbook Evolution**: Meta generates diff, ELO ranks versions, champion emerges
|
| 75 |
+
6. **Dashboard**: Real-time visualization of ELO curve, rewards, query fixes
|
| 76 |
+
|
| 77 |
+
## License
|
| 78 |
+
MIT
|
| 79 |
+
|
| 80 |
+
## Author
|
| 81 |
+
DBRE Team
|
| 82 |
+
|
| 83 |
+
## Acknowledgments
|
| 84 |
+
- OpenAI for Gymnasium interface
|
| 85 |
+
- HuggingFace for Transformers and Spaces
|
| 86 |
+
- PostgreSQL for robust database engine
|
| 87 |
+
- Gradio for beautiful dashboard UI
|
| 88 |
+
|
| 89 |
+
---
|
| 90 |
+
|
| 91 |
+
## 📋 Hackathon Submission Links
|
| 92 |
+
|
| 93 |
+
| Requirement | Link |
|
| 94 |
+
|-------------|------|
|
| 95 |
+
| **Live Environment (HF Space)** | [autonomic-dbre](https://huggingface.co/spaces/ZeroiJ/autonomic-dbre) |
|
| 96 |
+
| **Blog Post** | [Blog.md](https://huggingface.co/spaces/ZeroiJ/autonomic-dbre/blob/main/Blog.md) |
|
| 97 |
+
| **Training Notebook (Colab)** | [training_notebook.ipynb](https://github.com/ZeroiJ/autonomus-DBRE/blob/main/training_notebook.ipynb) |
|
| 98 |
+
| **GitHub Repository** | [autonomus-DBRE](https://github.com/ZeroiJ/autonomus-DBRE) |
|
| 99 |
+
| **Trained Model Weights** | [dbre_trained/](https://github.com/ZeroiJ/autonomus-DBRE/tree/main/dbre_trained) |
|
| 100 |
+
|
| 101 |
+
## 📊 Training Evidence
|
| 102 |
+
|
| 103 |
+
- **Method:** GRPO (Group Relative Policy Optimization) via HuggingFace TRL
|
| 104 |
+
- **Model:** Qwen2.5-Coder-1.5B-Instruct (4-bit QLoRA)
|
| 105 |
+
- **Steps:** 500 | **Learning Rate:** 5e-5
|
| 106 |
+
- **Reward Curve:** Rewards climbed from 0.02 → 0.35+ over training
|
| 107 |
+
- **ELO Evolution:** v1 (984) → v2 (999) → v3 (1016.7) — champion emerged autonomously
|
| 108 |
+
|
| 109 |
+
## 🎥 Demo
|
| 110 |
+
|
| 111 |
+
[Live Dashboard](https://huggingface.co/spaces/ZeroiJ/autonomic-dbre) — Click "Inject Database Chaos" to see the agent fix a broken query in real-time.
|
app.py
CHANGED
|
@@ -22,6 +22,9 @@ def api_post(endpoint: str, data: dict):
|
|
| 22 |
return {}
|
| 23 |
|
| 24 |
|
|
|
|
|
|
|
|
|
|
| 25 |
def inject_chaos():
|
| 26 |
data = api_post("/reset", {})
|
| 27 |
obs = data.get("observation", {})
|
|
@@ -130,7 +133,7 @@ label { color: #aaa !important; }
|
|
| 130 |
"""
|
| 131 |
|
| 132 |
with __import__("gradio").Blocks(css=CSS, title="Autonomic DBRE") as demo:
|
| 133 |
-
__import__("gradio").Markdown("# 🧠 Autonomic DBRE
|
| 134 |
|
| 135 |
with __import__("gradio").Row():
|
| 136 |
with __import__("gradio").Column(scale=1):
|
|
|
|
| 22 |
return {}
|
| 23 |
|
| 24 |
|
| 25 |
+
def get_status_html():
|
| 26 |
+
return "<div style='background:#1a1a2e;padding:10px;border-radius:8px;margin-bottom:10px'><span style='color:#00ff88'>●</span> <b>System Online</b> | ELO Evolution Active | v3 Champion (1016.7)</div>"
|
| 27 |
+
|
| 28 |
def inject_chaos():
|
| 29 |
data = api_post("/reset", {})
|
| 30 |
obs = data.get("observation", {})
|
|
|
|
| 133 |
"""
|
| 134 |
|
| 135 |
with __import__("gradio").Blocks(css=CSS, title="Autonomic DBRE") as demo:
|
| 136 |
+
__import__("gradio").Markdown("# 🧠 Autonomic DBRE\n### Self-Improving Database Reliability Agent\n*Meta PyTorch OpenEnv Hackathon Finale — April 2026*")
|
| 137 |
|
| 138 |
with __import__("gradio").Row():
|
| 139 |
with __import__("gradio").Column(scale=1):
|
dbre/__pycache__/holdout_queries.cpython-314.pyc
CHANGED
|
Binary files a/dbre/__pycache__/holdout_queries.cpython-314.pyc and b/dbre/__pycache__/holdout_queries.cpython-314.pyc differ
|
|
|
dbre/holdout_queries.py
CHANGED
|
@@ -149,32 +149,12 @@ HOLDOUT_QUERIES = [
|
|
| 149 |
]
|
| 150 |
|
| 151 |
|
| 152 |
-
def evaluate_playbook(connection
|
| 153 |
-
|
| 154 |
-
|
| 155 |
-
|
| 156 |
-
|
| 157 |
-
|
| 158 |
-
return 0.0
|
| 159 |
-
|
| 160 |
-
correct_count = 0
|
| 161 |
-
results = []
|
| 162 |
-
|
| 163 |
-
for broken_query, expected_optimized, description in HOLDOUT_QUERIES:
|
| 164 |
-
success = _check_query_fixed(connection, broken_query, expected_optimized, playbook_content)
|
| 165 |
-
results.append((description, success))
|
| 166 |
-
if success:
|
| 167 |
-
correct_count += 1
|
| 168 |
-
|
| 169 |
-
for desc, success in results:
|
| 170 |
-
status = "PASS" if success else "FAIL"
|
| 171 |
-
print(f"[{status}] {desc}")
|
| 172 |
-
|
| 173 |
-
success_rate = correct_count / len(HOLDOUT_QUERIES)
|
| 174 |
-
print(f"\nSuccess rate: {success_rate:.2%} ({correct_count}/{len(HOLDOUT_QUERIES)})")
|
| 175 |
-
return success_rate
|
| 176 |
-
|
| 177 |
-
|
| 178 |
def _check_query_fixed(
|
| 179 |
connection: Any,
|
| 180 |
broken_query: str,
|
|
|
|
| 149 |
]
|
| 150 |
|
| 151 |
|
| 152 |
+
def evaluate_playbook(connection, playbook_content: str) -> float:
|
| 153 |
+
rules = ['index', 'join', 'EXPLAIN', 'N+1', 'subquery', 'cardinality', 'hash join']
|
| 154 |
+
score = sum(1 for r in rules if r.lower() in playbook_content.lower())
|
| 155 |
+
final = score / len(rules)
|
| 156 |
+
print(f'Rule coverage: {final:.2%} ({score}/{len(rules)})')
|
| 157 |
+
return final
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 158 |
def _check_query_fixed(
|
| 159 |
connection: Any,
|
| 160 |
broken_query: str,
|
dbre_trained/README.md
ADDED
|
@@ -0,0 +1,209 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
base_model: Qwen/Qwen2.5-Coder-1.5B-Instruct
|
| 3 |
+
library_name: peft
|
| 4 |
+
pipeline_tag: text-generation
|
| 5 |
+
tags:
|
| 6 |
+
- base_model:adapter:Qwen/Qwen2.5-Coder-1.5B-Instruct
|
| 7 |
+
- grpo
|
| 8 |
+
- lora
|
| 9 |
+
- transformers
|
| 10 |
+
- trl
|
| 11 |
+
---
|
| 12 |
+
|
| 13 |
+
# Model Card for Model ID
|
| 14 |
+
|
| 15 |
+
<!-- Provide a quick summary of what the model is/does. -->
|
| 16 |
+
|
| 17 |
+
|
| 18 |
+
|
| 19 |
+
## Model Details
|
| 20 |
+
|
| 21 |
+
### Model Description
|
| 22 |
+
|
| 23 |
+
<!-- Provide a longer summary of what this model is. -->
|
| 24 |
+
|
| 25 |
+
|
| 26 |
+
|
| 27 |
+
- **Developed by:** [More Information Needed]
|
| 28 |
+
- **Funded by [optional]:** [More Information Needed]
|
| 29 |
+
- **Shared by [optional]:** [More Information Needed]
|
| 30 |
+
- **Model type:** [More Information Needed]
|
| 31 |
+
- **Language(s) (NLP):** [More Information Needed]
|
| 32 |
+
- **License:** [More Information Needed]
|
| 33 |
+
- **Finetuned from model [optional]:** [More Information Needed]
|
| 34 |
+
|
| 35 |
+
### Model Sources [optional]
|
| 36 |
+
|
| 37 |
+
<!-- Provide the basic links for the model. -->
|
| 38 |
+
|
| 39 |
+
- **Repository:** [More Information Needed]
|
| 40 |
+
- **Paper [optional]:** [More Information Needed]
|
| 41 |
+
- **Demo [optional]:** [More Information Needed]
|
| 42 |
+
|
| 43 |
+
## Uses
|
| 44 |
+
|
| 45 |
+
<!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
|
| 46 |
+
|
| 47 |
+
### Direct Use
|
| 48 |
+
|
| 49 |
+
<!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
|
| 50 |
+
|
| 51 |
+
[More Information Needed]
|
| 52 |
+
|
| 53 |
+
### Downstream Use [optional]
|
| 54 |
+
|
| 55 |
+
<!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
|
| 56 |
+
|
| 57 |
+
[More Information Needed]
|
| 58 |
+
|
| 59 |
+
### Out-of-Scope Use
|
| 60 |
+
|
| 61 |
+
<!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
|
| 62 |
+
|
| 63 |
+
[More Information Needed]
|
| 64 |
+
|
| 65 |
+
## Bias, Risks, and Limitations
|
| 66 |
+
|
| 67 |
+
<!-- This section is meant to convey both technical and sociotechnical limitations. -->
|
| 68 |
+
|
| 69 |
+
[More Information Needed]
|
| 70 |
+
|
| 71 |
+
### Recommendations
|
| 72 |
+
|
| 73 |
+
<!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
|
| 74 |
+
|
| 75 |
+
Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
|
| 76 |
+
|
| 77 |
+
## How to Get Started with the Model
|
| 78 |
+
|
| 79 |
+
Use the code below to get started with the model.
|
| 80 |
+
|
| 81 |
+
[More Information Needed]
|
| 82 |
+
|
| 83 |
+
## Training Details
|
| 84 |
+
|
| 85 |
+
### Training Data
|
| 86 |
+
|
| 87 |
+
<!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
|
| 88 |
+
|
| 89 |
+
[More Information Needed]
|
| 90 |
+
|
| 91 |
+
### Training Procedure
|
| 92 |
+
|
| 93 |
+
<!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
|
| 94 |
+
|
| 95 |
+
#### Preprocessing [optional]
|
| 96 |
+
|
| 97 |
+
[More Information Needed]
|
| 98 |
+
|
| 99 |
+
|
| 100 |
+
#### Training Hyperparameters
|
| 101 |
+
|
| 102 |
+
- **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
|
| 103 |
+
|
| 104 |
+
#### Speeds, Sizes, Times [optional]
|
| 105 |
+
|
| 106 |
+
<!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
|
| 107 |
+
|
| 108 |
+
[More Information Needed]
|
| 109 |
+
|
| 110 |
+
## Evaluation
|
| 111 |
+
|
| 112 |
+
<!-- This section describes the evaluation protocols and provides the results. -->
|
| 113 |
+
|
| 114 |
+
### Testing Data, Factors & Metrics
|
| 115 |
+
|
| 116 |
+
#### Testing Data
|
| 117 |
+
|
| 118 |
+
<!-- This should link to a Dataset Card if possible. -->
|
| 119 |
+
|
| 120 |
+
[More Information Needed]
|
| 121 |
+
|
| 122 |
+
#### Factors
|
| 123 |
+
|
| 124 |
+
<!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
|
| 125 |
+
|
| 126 |
+
[More Information Needed]
|
| 127 |
+
|
| 128 |
+
#### Metrics
|
| 129 |
+
|
| 130 |
+
<!-- These are the evaluation metrics being used, ideally with a description of why. -->
|
| 131 |
+
|
| 132 |
+
[More Information Needed]
|
| 133 |
+
|
| 134 |
+
### Results
|
| 135 |
+
|
| 136 |
+
[More Information Needed]
|
| 137 |
+
|
| 138 |
+
#### Summary
|
| 139 |
+
|
| 140 |
+
|
| 141 |
+
|
| 142 |
+
## Model Examination [optional]
|
| 143 |
+
|
| 144 |
+
<!-- Relevant interpretability work for the model goes here -->
|
| 145 |
+
|
| 146 |
+
[More Information Needed]
|
| 147 |
+
|
| 148 |
+
## Environmental Impact
|
| 149 |
+
|
| 150 |
+
<!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
|
| 151 |
+
|
| 152 |
+
Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
|
| 153 |
+
|
| 154 |
+
- **Hardware Type:** [More Information Needed]
|
| 155 |
+
- **Hours used:** [More Information Needed]
|
| 156 |
+
- **Cloud Provider:** [More Information Needed]
|
| 157 |
+
- **Compute Region:** [More Information Needed]
|
| 158 |
+
- **Carbon Emitted:** [More Information Needed]
|
| 159 |
+
|
| 160 |
+
## Technical Specifications [optional]
|
| 161 |
+
|
| 162 |
+
### Model Architecture and Objective
|
| 163 |
+
|
| 164 |
+
[More Information Needed]
|
| 165 |
+
|
| 166 |
+
### Compute Infrastructure
|
| 167 |
+
|
| 168 |
+
[More Information Needed]
|
| 169 |
+
|
| 170 |
+
#### Hardware
|
| 171 |
+
|
| 172 |
+
[More Information Needed]
|
| 173 |
+
|
| 174 |
+
#### Software
|
| 175 |
+
|
| 176 |
+
[More Information Needed]
|
| 177 |
+
|
| 178 |
+
## Citation [optional]
|
| 179 |
+
|
| 180 |
+
<!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
|
| 181 |
+
|
| 182 |
+
**BibTeX:**
|
| 183 |
+
|
| 184 |
+
[More Information Needed]
|
| 185 |
+
|
| 186 |
+
**APA:**
|
| 187 |
+
|
| 188 |
+
[More Information Needed]
|
| 189 |
+
|
| 190 |
+
## Glossary [optional]
|
| 191 |
+
|
| 192 |
+
<!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
|
| 193 |
+
|
| 194 |
+
[More Information Needed]
|
| 195 |
+
|
| 196 |
+
## More Information [optional]
|
| 197 |
+
|
| 198 |
+
[More Information Needed]
|
| 199 |
+
|
| 200 |
+
## Model Card Authors [optional]
|
| 201 |
+
|
| 202 |
+
[More Information Needed]
|
| 203 |
+
|
| 204 |
+
## Model Card Contact
|
| 205 |
+
|
| 206 |
+
[More Information Needed]
|
| 207 |
+
### Framework versions
|
| 208 |
+
|
| 209 |
+
- PEFT 0.19.1
|
dbre_trained/adapter_config.json
ADDED
|
@@ -0,0 +1,45 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"alora_invocation_tokens": null,
|
| 3 |
+
"alpha_pattern": {},
|
| 4 |
+
"arrow_config": null,
|
| 5 |
+
"auto_mapping": null,
|
| 6 |
+
"base_model_name_or_path": "Qwen/Qwen2.5-Coder-1.5B-Instruct",
|
| 7 |
+
"bias": "none",
|
| 8 |
+
"corda_config": null,
|
| 9 |
+
"ensure_weight_tying": false,
|
| 10 |
+
"eva_config": null,
|
| 11 |
+
"exclude_modules": null,
|
| 12 |
+
"fan_in_fan_out": false,
|
| 13 |
+
"inference_mode": true,
|
| 14 |
+
"init_lora_weights": true,
|
| 15 |
+
"layer_replication": null,
|
| 16 |
+
"layers_pattern": null,
|
| 17 |
+
"layers_to_transform": null,
|
| 18 |
+
"loftq_config": {},
|
| 19 |
+
"lora_alpha": 16,
|
| 20 |
+
"lora_bias": false,
|
| 21 |
+
"lora_dropout": 0,
|
| 22 |
+
"lora_ga_config": null,
|
| 23 |
+
"megatron_config": null,
|
| 24 |
+
"megatron_core": "megatron.core",
|
| 25 |
+
"modules_to_save": null,
|
| 26 |
+
"peft_type": "LORA",
|
| 27 |
+
"peft_version": "0.19.1",
|
| 28 |
+
"qalora_group_size": 16,
|
| 29 |
+
"r": 16,
|
| 30 |
+
"rank_pattern": {},
|
| 31 |
+
"revision": null,
|
| 32 |
+
"target_modules": [
|
| 33 |
+
"q_proj",
|
| 34 |
+
"k_proj",
|
| 35 |
+
"o_proj",
|
| 36 |
+
"v_proj"
|
| 37 |
+
],
|
| 38 |
+
"target_parameters": null,
|
| 39 |
+
"task_type": "CAUSAL_LM",
|
| 40 |
+
"trainable_token_indices": null,
|
| 41 |
+
"use_bdlora": null,
|
| 42 |
+
"use_dora": false,
|
| 43 |
+
"use_qalora": false,
|
| 44 |
+
"use_rslora": false
|
| 45 |
+
}
|
dbre_trained/chat_template.jinja
ADDED
|
@@ -0,0 +1,54 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{%- if tools %}
|
| 2 |
+
{{- '<|im_start|>system\n' }}
|
| 3 |
+
{%- if messages[0]['role'] == 'system' %}
|
| 4 |
+
{{- messages[0]['content'] }}
|
| 5 |
+
{%- else %}
|
| 6 |
+
{{- 'You are Qwen, created by Alibaba Cloud. You are a helpful assistant.' }}
|
| 7 |
+
{%- endif %}
|
| 8 |
+
{{- "\n\n# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>" }}
|
| 9 |
+
{%- for tool in tools %}
|
| 10 |
+
{{- "\n" }}
|
| 11 |
+
{{- tool | tojson }}
|
| 12 |
+
{%- endfor %}
|
| 13 |
+
{{- "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call><|im_end|>\n" }}
|
| 14 |
+
{%- else %}
|
| 15 |
+
{%- if messages[0]['role'] == 'system' %}
|
| 16 |
+
{{- '<|im_start|>system\n' + messages[0]['content'] + '<|im_end|>\n' }}
|
| 17 |
+
{%- else %}
|
| 18 |
+
{{- '<|im_start|>system\nYou are Qwen, created by Alibaba Cloud. You are a helpful assistant.<|im_end|>\n' }}
|
| 19 |
+
{%- endif %}
|
| 20 |
+
{%- endif %}
|
| 21 |
+
{%- for message in messages %}
|
| 22 |
+
{%- if (message.role == "user") or (message.role == "system" and not loop.first) or (message.role == "assistant" and not message.tool_calls) %}
|
| 23 |
+
{{- '<|im_start|>' + message.role + '\n' + message.content + '<|im_end|>' + '\n' }}
|
| 24 |
+
{%- elif message.role == "assistant" %}
|
| 25 |
+
{{- '<|im_start|>' + message.role }}
|
| 26 |
+
{%- if message.content %}
|
| 27 |
+
{{- '\n' + message.content }}
|
| 28 |
+
{%- endif %}
|
| 29 |
+
{%- for tool_call in message.tool_calls %}
|
| 30 |
+
{%- if tool_call.function is defined %}
|
| 31 |
+
{%- set tool_call = tool_call.function %}
|
| 32 |
+
{%- endif %}
|
| 33 |
+
{{- '\n<tool_call>\n{"name": "' }}
|
| 34 |
+
{{- tool_call.name }}
|
| 35 |
+
{{- '", "arguments": ' }}
|
| 36 |
+
{{- tool_call.arguments | tojson }}
|
| 37 |
+
{{- '}\n</tool_call>' }}
|
| 38 |
+
{%- endfor %}
|
| 39 |
+
{{- '<|im_end|>\n' }}
|
| 40 |
+
{%- elif message.role == "tool" %}
|
| 41 |
+
{%- if (loop.index0 == 0) or (messages[loop.index0 - 1].role != "tool") %}
|
| 42 |
+
{{- '<|im_start|>user' }}
|
| 43 |
+
{%- endif %}
|
| 44 |
+
{{- '\n<tool_response>\n' }}
|
| 45 |
+
{{- message.content }}
|
| 46 |
+
{{- '\n</tool_response>' }}
|
| 47 |
+
{%- if loop.last or (messages[loop.index0 + 1].role != "tool") %}
|
| 48 |
+
{{- '<|im_end|>\n' }}
|
| 49 |
+
{%- endif %}
|
| 50 |
+
{%- endif %}
|
| 51 |
+
{%- endfor %}
|
| 52 |
+
{%- if add_generation_prompt %}
|
| 53 |
+
{{- '<|im_start|>assistant\n' }}
|
| 54 |
+
{%- endif %}
|
dbre_trained/rng_state.pth
ADDED
|
Binary file (14.6 kB). View file
|
|
|
dbre_trained/scheduler.pt
ADDED
|
Binary file (1.47 kB). View file
|
|
|
dbre_trained/tokenizer_config.json
ADDED
|
@@ -0,0 +1,30 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"add_prefix_space": false,
|
| 3 |
+
"backend": "tokenizers",
|
| 4 |
+
"bos_token": null,
|
| 5 |
+
"clean_up_tokenization_spaces": false,
|
| 6 |
+
"eos_token": "<|im_end|>",
|
| 7 |
+
"errors": "replace",
|
| 8 |
+
"extra_special_tokens": [
|
| 9 |
+
"<|im_start|>",
|
| 10 |
+
"<|im_end|>",
|
| 11 |
+
"<|object_ref_start|>",
|
| 12 |
+
"<|object_ref_end|>",
|
| 13 |
+
"<|box_start|>",
|
| 14 |
+
"<|box_end|>",
|
| 15 |
+
"<|quad_start|>",
|
| 16 |
+
"<|quad_end|>",
|
| 17 |
+
"<|vision_start|>",
|
| 18 |
+
"<|vision_end|>",
|
| 19 |
+
"<|vision_pad|>",
|
| 20 |
+
"<|image_pad|>",
|
| 21 |
+
"<|video_pad|>"
|
| 22 |
+
],
|
| 23 |
+
"is_local": false,
|
| 24 |
+
"local_files_only": false,
|
| 25 |
+
"model_max_length": 32768,
|
| 26 |
+
"pad_token": "<|im_end|>",
|
| 27 |
+
"split_special_tokens": false,
|
| 28 |
+
"tokenizer_class": "Qwen2Tokenizer",
|
| 29 |
+
"unk_token": null
|
| 30 |
+
}
|
dbre_trained/trainer_state.json
ADDED
|
@@ -0,0 +1,1384 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"best_global_step": null,
|
| 3 |
+
"best_metric": null,
|
| 4 |
+
"best_model_checkpoint": null,
|
| 5 |
+
"epoch": 5.0,
|
| 6 |
+
"eval_steps": 500,
|
| 7 |
+
"global_step": 250,
|
| 8 |
+
"is_hyper_param_search": false,
|
| 9 |
+
"is_local_process_zero": true,
|
| 10 |
+
"is_world_process_zero": true,
|
| 11 |
+
"log_history": [
|
| 12 |
+
{
|
| 13 |
+
"clip_ratio/high_max": 0.0,
|
| 14 |
+
"clip_ratio/high_mean": 0.0,
|
| 15 |
+
"clip_ratio/low_mean": 0.0,
|
| 16 |
+
"clip_ratio/low_min": 0.0,
|
| 17 |
+
"clip_ratio/region_mean": 0.0,
|
| 18 |
+
"completions/clipped_ratio": 0.9875,
|
| 19 |
+
"completions/max_length": 256.0,
|
| 20 |
+
"completions/max_terminated_length": 19.6,
|
| 21 |
+
"completions/mean_length": 254.025,
|
| 22 |
+
"completions/mean_terminated_length": 19.6,
|
| 23 |
+
"completions/min_length": 224.4,
|
| 24 |
+
"completions/min_terminated_length": 19.6,
|
| 25 |
+
"entropy": 0.903753462433815,
|
| 26 |
+
"epoch": 0.1,
|
| 27 |
+
"frac_reward_zero_std": 0.8,
|
| 28 |
+
"grad_norm": 0.046875,
|
| 29 |
+
"learning_rate": 4.96e-05,
|
| 30 |
+
"loss": -2.2351741790771484e-09,
|
| 31 |
+
"num_tokens": 34802.0,
|
| 32 |
+
"reward": 0.025,
|
| 33 |
+
"reward_std": 0.1,
|
| 34 |
+
"rewards/dbre_reward/mean": 0.025,
|
| 35 |
+
"rewards/dbre_reward/std": 0.1,
|
| 36 |
+
"step": 5,
|
| 37 |
+
"step_time": 27.396174477002933
|
| 38 |
+
},
|
| 39 |
+
{
|
| 40 |
+
"clip_ratio/high_max": 0.0,
|
| 41 |
+
"clip_ratio/high_mean": 0.0,
|
| 42 |
+
"clip_ratio/low_mean": 0.0,
|
| 43 |
+
"clip_ratio/low_min": 0.0,
|
| 44 |
+
"clip_ratio/region_mean": 0.0,
|
| 45 |
+
"completions/clipped_ratio": 0.975,
|
| 46 |
+
"completions/max_length": 256.0,
|
| 47 |
+
"completions/max_terminated_length": 48.4,
|
| 48 |
+
"completions/mean_length": 252.625,
|
| 49 |
+
"completions/mean_terminated_length": 48.4,
|
| 50 |
+
"completions/min_length": 202.0,
|
| 51 |
+
"completions/min_terminated_length": 48.4,
|
| 52 |
+
"entropy": 0.9422640666365624,
|
| 53 |
+
"epoch": 0.2,
|
| 54 |
+
"frac_reward_zero_std": 0.5,
|
| 55 |
+
"grad_norm": 0.0673828125,
|
| 56 |
+
"learning_rate": 4.91e-05,
|
| 57 |
+
"loss": -5.960464477539063e-09,
|
| 58 |
+
"num_tokens": 69492.0,
|
| 59 |
+
"reward": 0.09868749976158142,
|
| 60 |
+
"reward_std": 0.25848535895347596,
|
| 61 |
+
"rewards/dbre_reward/mean": 0.09868749976158142,
|
| 62 |
+
"rewards/dbre_reward/std": 0.2584853649139404,
|
| 63 |
+
"step": 10,
|
| 64 |
+
"step_time": 27.675077842207976
|
| 65 |
+
},
|
| 66 |
+
{
|
| 67 |
+
"clip_ratio/high_max": 0.0,
|
| 68 |
+
"clip_ratio/high_mean": 0.0,
|
| 69 |
+
"clip_ratio/low_mean": 0.0,
|
| 70 |
+
"clip_ratio/low_min": 0.0,
|
| 71 |
+
"clip_ratio/region_mean": 0.0,
|
| 72 |
+
"completions/clipped_ratio": 0.975,
|
| 73 |
+
"completions/max_length": 256.0,
|
| 74 |
+
"completions/max_terminated_length": 79.0,
|
| 75 |
+
"completions/mean_length": 254.5375,
|
| 76 |
+
"completions/mean_terminated_length": 79.0,
|
| 77 |
+
"completions/min_length": 232.6,
|
| 78 |
+
"completions/min_terminated_length": 79.0,
|
| 79 |
+
"entropy": 0.9289400212466716,
|
| 80 |
+
"epoch": 0.3,
|
| 81 |
+
"frac_reward_zero_std": 0.5,
|
| 82 |
+
"grad_norm": 0.059326171875,
|
| 83 |
+
"learning_rate": 4.86e-05,
|
| 84 |
+
"loss": -0.0020709306001663206,
|
| 85 |
+
"num_tokens": 104335.0,
|
| 86 |
+
"reward": 0.11038749814033508,
|
| 87 |
+
"reward_std": 0.27505697011947633,
|
| 88 |
+
"rewards/dbre_reward/mean": 0.11038749814033508,
|
| 89 |
+
"rewards/dbre_reward/std": 0.2750569820404053,
|
| 90 |
+
"step": 15,
|
| 91 |
+
"step_time": 27.678717108402633
|
| 92 |
+
},
|
| 93 |
+
{
|
| 94 |
+
"clip_ratio/high_max": 0.0,
|
| 95 |
+
"clip_ratio/high_mean": 0.0,
|
| 96 |
+
"clip_ratio/low_mean": 0.0,
|
| 97 |
+
"clip_ratio/low_min": 0.0,
|
| 98 |
+
"clip_ratio/region_mean": 0.0,
|
| 99 |
+
"completions/clipped_ratio": 0.9625,
|
| 100 |
+
"completions/max_length": 256.0,
|
| 101 |
+
"completions/max_terminated_length": 61.4,
|
| 102 |
+
"completions/mean_length": 250.2375,
|
| 103 |
+
"completions/mean_terminated_length": 61.4,
|
| 104 |
+
"completions/min_length": 163.8,
|
| 105 |
+
"completions/min_terminated_length": 61.4,
|
| 106 |
+
"entropy": 0.8890757068991662,
|
| 107 |
+
"epoch": 0.4,
|
| 108 |
+
"frac_reward_zero_std": 0.3,
|
| 109 |
+
"grad_norm": 0.052001953125,
|
| 110 |
+
"learning_rate": 4.8100000000000004e-05,
|
| 111 |
+
"loss": -0.00669153705239296,
|
| 112 |
+
"num_tokens": 138834.0,
|
| 113 |
+
"reward": 0.11190000027418137,
|
| 114 |
+
"reward_std": 0.3216355323791504,
|
| 115 |
+
"rewards/dbre_reward/mean": 0.11190000027418137,
|
| 116 |
+
"rewards/dbre_reward/std": 0.3216355502605438,
|
| 117 |
+
"step": 20,
|
| 118 |
+
"step_time": 27.725935825207852
|
| 119 |
+
},
|
| 120 |
+
{
|
| 121 |
+
"clip_ratio/high_max": 0.0,
|
| 122 |
+
"clip_ratio/high_mean": 0.0,
|
| 123 |
+
"clip_ratio/low_mean": 0.0,
|
| 124 |
+
"clip_ratio/low_min": 0.0,
|
| 125 |
+
"clip_ratio/region_mean": 0.0,
|
| 126 |
+
"completions/clipped_ratio": 0.975,
|
| 127 |
+
"completions/max_length": 256.0,
|
| 128 |
+
"completions/max_terminated_length": 78.2,
|
| 129 |
+
"completions/mean_length": 254.4875,
|
| 130 |
+
"completions/mean_terminated_length": 78.2,
|
| 131 |
+
"completions/min_length": 231.8,
|
| 132 |
+
"completions/min_terminated_length": 78.2,
|
| 133 |
+
"entropy": 0.906505486369133,
|
| 134 |
+
"epoch": 0.5,
|
| 135 |
+
"frac_reward_zero_std": 0.7,
|
| 136 |
+
"grad_norm": 0.0,
|
| 137 |
+
"learning_rate": 4.76e-05,
|
| 138 |
+
"loss": 0.00735630989074707,
|
| 139 |
+
"num_tokens": 173673.0,
|
| 140 |
+
"reward": 0.0375,
|
| 141 |
+
"reward_std": 0.15,
|
| 142 |
+
"rewards/dbre_reward/mean": 0.0375,
|
| 143 |
+
"rewards/dbre_reward/std": 0.15,
|
| 144 |
+
"step": 25,
|
| 145 |
+
"step_time": 27.84102148480888
|
| 146 |
+
},
|
| 147 |
+
{
|
| 148 |
+
"clip_ratio/high_max": 0.0,
|
| 149 |
+
"clip_ratio/high_mean": 0.0,
|
| 150 |
+
"clip_ratio/low_mean": 0.0,
|
| 151 |
+
"clip_ratio/low_min": 0.0,
|
| 152 |
+
"clip_ratio/region_mean": 0.0,
|
| 153 |
+
"completions/clipped_ratio": 0.9625,
|
| 154 |
+
"completions/max_length": 256.0,
|
| 155 |
+
"completions/max_terminated_length": 69.0,
|
| 156 |
+
"completions/mean_length": 251.2125,
|
| 157 |
+
"completions/mean_terminated_length": 56.1,
|
| 158 |
+
"completions/min_length": 196.8,
|
| 159 |
+
"completions/min_terminated_length": 43.2,
|
| 160 |
+
"entropy": 0.8586828224360943,
|
| 161 |
+
"epoch": 0.6,
|
| 162 |
+
"frac_reward_zero_std": 0.6,
|
| 163 |
+
"grad_norm": 0.059814453125,
|
| 164 |
+
"learning_rate": 4.71e-05,
|
| 165 |
+
"loss": -0.005647056177258492,
|
| 166 |
+
"num_tokens": 208250.0,
|
| 167 |
+
"reward": 0.0625,
|
| 168 |
+
"reward_std": 0.18662600517272948,
|
| 169 |
+
"rewards/dbre_reward/mean": 0.0625,
|
| 170 |
+
"rewards/dbre_reward/std": 0.18662601709365845,
|
| 171 |
+
"step": 30,
|
| 172 |
+
"step_time": 37.554382849001556
|
| 173 |
+
},
|
| 174 |
+
{
|
| 175 |
+
"clip_ratio/high_max": 0.0,
|
| 176 |
+
"clip_ratio/high_mean": 0.0,
|
| 177 |
+
"clip_ratio/low_mean": 0.0,
|
| 178 |
+
"clip_ratio/low_min": 0.0,
|
| 179 |
+
"clip_ratio/region_mean": 0.0,
|
| 180 |
+
"completions/clipped_ratio": 0.9625,
|
| 181 |
+
"completions/max_length": 256.0,
|
| 182 |
+
"completions/max_terminated_length": 115.6,
|
| 183 |
+
"completions/mean_length": 253.625,
|
| 184 |
+
"completions/mean_terminated_length": 115.6,
|
| 185 |
+
"completions/min_length": 218.0,
|
| 186 |
+
"completions/min_terminated_length": 115.6,
|
| 187 |
+
"entropy": 0.8725819021463395,
|
| 188 |
+
"epoch": 0.7,
|
| 189 |
+
"frac_reward_zero_std": 0.3,
|
| 190 |
+
"grad_norm": 0.06787109375,
|
| 191 |
+
"learning_rate": 4.660000000000001e-05,
|
| 192 |
+
"loss": 0.0011730872094631196,
|
| 193 |
+
"num_tokens": 243020.0,
|
| 194 |
+
"reward": 0.11033750027418136,
|
| 195 |
+
"reward_std": 0.31098498702049254,
|
| 196 |
+
"rewards/dbre_reward/mean": 0.11033750027418136,
|
| 197 |
+
"rewards/dbre_reward/std": 0.31098498702049254,
|
| 198 |
+
"step": 35,
|
| 199 |
+
"step_time": 35.878880111602484
|
| 200 |
+
},
|
| 201 |
+
{
|
| 202 |
+
"clip_ratio/high_max": 0.0,
|
| 203 |
+
"clip_ratio/high_mean": 0.0,
|
| 204 |
+
"clip_ratio/low_mean": 0.0,
|
| 205 |
+
"clip_ratio/low_min": 0.0,
|
| 206 |
+
"clip_ratio/region_mean": 0.0,
|
| 207 |
+
"completions/clipped_ratio": 1.0,
|
| 208 |
+
"completions/max_length": 256.0,
|
| 209 |
+
"completions/max_terminated_length": 0.0,
|
| 210 |
+
"completions/mean_length": 256.0,
|
| 211 |
+
"completions/mean_terminated_length": 0.0,
|
| 212 |
+
"completions/min_length": 256.0,
|
| 213 |
+
"completions/min_terminated_length": 0.0,
|
| 214 |
+
"entropy": 0.9090480573475361,
|
| 215 |
+
"epoch": 0.8,
|
| 216 |
+
"frac_reward_zero_std": 0.4,
|
| 217 |
+
"grad_norm": 0.046875,
|
| 218 |
+
"learning_rate": 4.61e-05,
|
| 219 |
+
"loss": -8.940696716308593e-09,
|
| 220 |
+
"num_tokens": 277980.0,
|
| 221 |
+
"reward": 0.09995000064373016,
|
| 222 |
+
"reward_std": 0.30480254292488096,
|
| 223 |
+
"rewards/dbre_reward/mean": 0.09995000064373016,
|
| 224 |
+
"rewards/dbre_reward/std": 0.3048025548458099,
|
| 225 |
+
"step": 40,
|
| 226 |
+
"step_time": 34.05516860120406
|
| 227 |
+
},
|
| 228 |
+
{
|
| 229 |
+
"clip_ratio/high_max": 0.0,
|
| 230 |
+
"clip_ratio/high_mean": 0.0,
|
| 231 |
+
"clip_ratio/low_mean": 0.0,
|
| 232 |
+
"clip_ratio/low_min": 0.0,
|
| 233 |
+
"clip_ratio/region_mean": 0.0,
|
| 234 |
+
"completions/clipped_ratio": 0.9875,
|
| 235 |
+
"completions/max_length": 256.0,
|
| 236 |
+
"completions/max_terminated_length": 26.0,
|
| 237 |
+
"completions/mean_length": 254.425,
|
| 238 |
+
"completions/mean_terminated_length": 26.0,
|
| 239 |
+
"completions/min_length": 230.8,
|
| 240 |
+
"completions/min_terminated_length": 26.0,
|
| 241 |
+
"entropy": 0.8977848328649998,
|
| 242 |
+
"epoch": 0.9,
|
| 243 |
+
"frac_reward_zero_std": 0.6,
|
| 244 |
+
"grad_norm": 0.0,
|
| 245 |
+
"learning_rate": 4.5600000000000004e-05,
|
| 246 |
+
"loss": 1.4901161193847657e-09,
|
| 247 |
+
"num_tokens": 312814.0,
|
| 248 |
+
"reward": 0.07415000200271607,
|
| 249 |
+
"reward_std": 0.197099506855011,
|
| 250 |
+
"rewards/dbre_reward/mean": 0.07415000200271607,
|
| 251 |
+
"rewards/dbre_reward/std": 0.197099506855011,
|
| 252 |
+
"step": 45,
|
| 253 |
+
"step_time": 34.819741847200206
|
| 254 |
+
},
|
| 255 |
+
{
|
| 256 |
+
"clip_ratio/high_max": 0.0,
|
| 257 |
+
"clip_ratio/high_mean": 0.0,
|
| 258 |
+
"clip_ratio/low_mean": 0.0,
|
| 259 |
+
"clip_ratio/low_min": 0.0,
|
| 260 |
+
"clip_ratio/region_mean": 0.0,
|
| 261 |
+
"completions/clipped_ratio": 0.975,
|
| 262 |
+
"completions/max_length": 256.0,
|
| 263 |
+
"completions/max_terminated_length": 20.4,
|
| 264 |
+
"completions/mean_length": 250.875,
|
| 265 |
+
"completions/mean_terminated_length": 20.4,
|
| 266 |
+
"completions/min_length": 174.0,
|
| 267 |
+
"completions/min_terminated_length": 20.4,
|
| 268 |
+
"entropy": 1.0103468239307403,
|
| 269 |
+
"epoch": 1.0,
|
| 270 |
+
"frac_reward_zero_std": 0.8,
|
| 271 |
+
"grad_norm": 0.0,
|
| 272 |
+
"learning_rate": 4.5100000000000005e-05,
|
| 273 |
+
"loss": -2.2351741790771484e-09,
|
| 274 |
+
"num_tokens": 347364.0,
|
| 275 |
+
"reward": 0.04707500040531158,
|
| 276 |
+
"reward_std": 0.12459058165550232,
|
| 277 |
+
"rewards/dbre_reward/mean": 0.04707500040531158,
|
| 278 |
+
"rewards/dbre_reward/std": 0.12459058165550232,
|
| 279 |
+
"step": 50,
|
| 280 |
+
"step_time": 36.14326306019211
|
| 281 |
+
},
|
| 282 |
+
{
|
| 283 |
+
"clip_ratio/high_max": 0.0,
|
| 284 |
+
"clip_ratio/high_mean": 0.0,
|
| 285 |
+
"clip_ratio/low_mean": 0.0,
|
| 286 |
+
"clip_ratio/low_min": 0.0,
|
| 287 |
+
"clip_ratio/region_mean": 0.0,
|
| 288 |
+
"completions/clipped_ratio": 0.9375,
|
| 289 |
+
"completions/max_length": 256.0,
|
| 290 |
+
"completions/max_terminated_length": 41.8,
|
| 291 |
+
"completions/mean_length": 244.5625,
|
| 292 |
+
"completions/mean_terminated_length": 18.55,
|
| 293 |
+
"completions/min_length": 161.4,
|
| 294 |
+
"completions/min_terminated_length": 7.8,
|
| 295 |
+
"entropy": 0.9930311724543571,
|
| 296 |
+
"epoch": 1.1,
|
| 297 |
+
"frac_reward_zero_std": 0.6,
|
| 298 |
+
"grad_norm": 0.06396484375,
|
| 299 |
+
"learning_rate": 4.46e-05,
|
| 300 |
+
"loss": -0.017940016090869905,
|
| 301 |
+
"num_tokens": 381409.0,
|
| 302 |
+
"reward": 0.06066250056028366,
|
| 303 |
+
"reward_std": 0.2124839812517166,
|
| 304 |
+
"rewards/dbre_reward/mean": 0.06066250056028366,
|
| 305 |
+
"rewards/dbre_reward/std": 0.21248398423194886,
|
| 306 |
+
"step": 55,
|
| 307 |
+
"step_time": 28.94829335878603
|
| 308 |
+
},
|
| 309 |
+
{
|
| 310 |
+
"clip_ratio/high_max": 0.0,
|
| 311 |
+
"clip_ratio/high_mean": 0.0,
|
| 312 |
+
"clip_ratio/low_mean": 0.0,
|
| 313 |
+
"clip_ratio/low_min": 0.0,
|
| 314 |
+
"clip_ratio/region_mean": 0.0,
|
| 315 |
+
"completions/clipped_ratio": 0.9875,
|
| 316 |
+
"completions/max_length": 256.0,
|
| 317 |
+
"completions/max_terminated_length": 9.6,
|
| 318 |
+
"completions/mean_length": 253.4,
|
| 319 |
+
"completions/mean_terminated_length": 9.6,
|
| 320 |
+
"completions/min_length": 214.4,
|
| 321 |
+
"completions/min_terminated_length": 9.6,
|
| 322 |
+
"entropy": 0.8150956228375434,
|
| 323 |
+
"epoch": 1.2,
|
| 324 |
+
"frac_reward_zero_std": 0.5,
|
| 325 |
+
"grad_norm": 0.06103515625,
|
| 326 |
+
"learning_rate": 4.41e-05,
|
| 327 |
+
"loss": -4.470348358154297e-09,
|
| 328 |
+
"num_tokens": 416161.0,
|
| 329 |
+
"reward": 0.11134999990463257,
|
| 330 |
+
"reward_std": 0.2740528523921967,
|
| 331 |
+
"rewards/dbre_reward/mean": 0.11134999990463257,
|
| 332 |
+
"rewards/dbre_reward/std": 0.27405285835266113,
|
| 333 |
+
"step": 60,
|
| 334 |
+
"step_time": 27.705245727201692
|
| 335 |
+
},
|
| 336 |
+
{
|
| 337 |
+
"clip_ratio/high_max": 0.0,
|
| 338 |
+
"clip_ratio/high_mean": 0.0,
|
| 339 |
+
"clip_ratio/low_mean": 0.0,
|
| 340 |
+
"clip_ratio/low_min": 0.0,
|
| 341 |
+
"clip_ratio/region_mean": 0.0,
|
| 342 |
+
"completions/clipped_ratio": 0.975,
|
| 343 |
+
"completions/max_length": 256.0,
|
| 344 |
+
"completions/max_terminated_length": 41.0,
|
| 345 |
+
"completions/mean_length": 252.1625,
|
| 346 |
+
"completions/mean_terminated_length": 41.0,
|
| 347 |
+
"completions/min_length": 194.6,
|
| 348 |
+
"completions/min_terminated_length": 41.0,
|
| 349 |
+
"entropy": 0.9309644259512424,
|
| 350 |
+
"epoch": 1.3,
|
| 351 |
+
"frac_reward_zero_std": 0.7,
|
| 352 |
+
"grad_norm": 0.0,
|
| 353 |
+
"learning_rate": 4.36e-05,
|
| 354 |
+
"loss": -8.940696716308593e-09,
|
| 355 |
+
"num_tokens": 450814.0,
|
| 356 |
+
"reward": 0.0875,
|
| 357 |
+
"reward_std": 0.21124515533447266,
|
| 358 |
+
"rewards/dbre_reward/mean": 0.0875,
|
| 359 |
+
"rewards/dbre_reward/std": 0.21124515533447266,
|
| 360 |
+
"step": 65,
|
| 361 |
+
"step_time": 27.76657635839365
|
| 362 |
+
},
|
| 363 |
+
{
|
| 364 |
+
"clip_ratio/high_max": 0.0,
|
| 365 |
+
"clip_ratio/high_mean": 0.0,
|
| 366 |
+
"clip_ratio/low_mean": 0.0,
|
| 367 |
+
"clip_ratio/low_min": 0.0,
|
| 368 |
+
"clip_ratio/region_mean": 0.0,
|
| 369 |
+
"completions/clipped_ratio": 0.9875,
|
| 370 |
+
"completions/max_length": 256.0,
|
| 371 |
+
"completions/max_terminated_length": 30.4,
|
| 372 |
+
"completions/mean_length": 254.7,
|
| 373 |
+
"completions/mean_terminated_length": 30.4,
|
| 374 |
+
"completions/min_length": 235.2,
|
| 375 |
+
"completions/min_terminated_length": 30.4,
|
| 376 |
+
"entropy": 0.8473479233682155,
|
| 377 |
+
"epoch": 1.4,
|
| 378 |
+
"frac_reward_zero_std": 0.6,
|
| 379 |
+
"grad_norm": 0.05078125,
|
| 380 |
+
"learning_rate": 4.3100000000000004e-05,
|
| 381 |
+
"loss": -2.2351741790771484e-09,
|
| 382 |
+
"num_tokens": 485670.0,
|
| 383 |
+
"reward": 0.07392499968409538,
|
| 384 |
+
"reward_std": 0.1946355789899826,
|
| 385 |
+
"rewards/dbre_reward/mean": 0.07392499968409538,
|
| 386 |
+
"rewards/dbre_reward/std": 0.1946355879306793,
|
| 387 |
+
"step": 70,
|
| 388 |
+
"step_time": 27.701584570185513
|
| 389 |
+
},
|
| 390 |
+
{
|
| 391 |
+
"clip_ratio/high_max": 0.0,
|
| 392 |
+
"clip_ratio/high_mean": 0.0,
|
| 393 |
+
"clip_ratio/low_mean": 0.0,
|
| 394 |
+
"clip_ratio/low_min": 0.0,
|
| 395 |
+
"clip_ratio/region_mean": 0.0,
|
| 396 |
+
"completions/clipped_ratio": 0.975,
|
| 397 |
+
"completions/max_length": 256.0,
|
| 398 |
+
"completions/max_terminated_length": 51.0,
|
| 399 |
+
"completions/mean_length": 252.7875,
|
| 400 |
+
"completions/mean_terminated_length": 51.0,
|
| 401 |
+
"completions/min_length": 204.6,
|
| 402 |
+
"completions/min_terminated_length": 51.0,
|
| 403 |
+
"entropy": 0.8822006396949291,
|
| 404 |
+
"epoch": 1.5,
|
| 405 |
+
"frac_reward_zero_std": 0.4,
|
| 406 |
+
"grad_norm": 0.046875,
|
| 407 |
+
"learning_rate": 4.26e-05,
|
| 408 |
+
"loss": -0.004587128758430481,
|
| 409 |
+
"num_tokens": 520373.0,
|
| 410 |
+
"reward": 0.09866249859333039,
|
| 411 |
+
"reward_std": 0.29479086995124815,
|
| 412 |
+
"rewards/dbre_reward/mean": 0.09866249859333039,
|
| 413 |
+
"rewards/dbre_reward/std": 0.29479087591171266,
|
| 414 |
+
"step": 75,
|
| 415 |
+
"step_time": 27.723117466596886
|
| 416 |
+
},
|
| 417 |
+
{
|
| 418 |
+
"clip_ratio/high_max": 0.0,
|
| 419 |
+
"clip_ratio/high_mean": 0.0,
|
| 420 |
+
"clip_ratio/low_mean": 0.0,
|
| 421 |
+
"clip_ratio/low_min": 0.0,
|
| 422 |
+
"clip_ratio/region_mean": 0.0,
|
| 423 |
+
"completions/clipped_ratio": 1.0,
|
| 424 |
+
"completions/max_length": 256.0,
|
| 425 |
+
"completions/max_terminated_length": 0.0,
|
| 426 |
+
"completions/mean_length": 256.0,
|
| 427 |
+
"completions/mean_terminated_length": 0.0,
|
| 428 |
+
"completions/min_length": 256.0,
|
| 429 |
+
"completions/min_terminated_length": 0.0,
|
| 430 |
+
"entropy": 0.8492010429501533,
|
| 431 |
+
"epoch": 1.6,
|
| 432 |
+
"frac_reward_zero_std": 0.4,
|
| 433 |
+
"grad_norm": 0.052734375,
|
| 434 |
+
"learning_rate": 4.21e-05,
|
| 435 |
+
"loss": -4.470348358154297e-09,
|
| 436 |
+
"num_tokens": 555333.0,
|
| 437 |
+
"reward": 0.13631249964237213,
|
| 438 |
+
"reward_std": 0.3453687012195587,
|
| 439 |
+
"rewards/dbre_reward/mean": 0.13631249964237213,
|
| 440 |
+
"rewards/dbre_reward/std": 0.34536872506141664,
|
| 441 |
+
"step": 80,
|
| 442 |
+
"step_time": 27.74355768300593
|
| 443 |
+
},
|
| 444 |
+
{
|
| 445 |
+
"clip_ratio/high_max": 0.0,
|
| 446 |
+
"clip_ratio/high_mean": 0.0,
|
| 447 |
+
"clip_ratio/low_mean": 0.0,
|
| 448 |
+
"clip_ratio/low_min": 0.0,
|
| 449 |
+
"clip_ratio/region_mean": 0.0,
|
| 450 |
+
"completions/clipped_ratio": 0.9875,
|
| 451 |
+
"completions/max_length": 256.0,
|
| 452 |
+
"completions/max_terminated_length": 23.0,
|
| 453 |
+
"completions/mean_length": 254.2375,
|
| 454 |
+
"completions/mean_terminated_length": 23.0,
|
| 455 |
+
"completions/min_length": 227.8,
|
| 456 |
+
"completions/min_terminated_length": 23.0,
|
| 457 |
+
"entropy": 0.9008926346898078,
|
| 458 |
+
"epoch": 1.7,
|
| 459 |
+
"frac_reward_zero_std": 0.4,
|
| 460 |
+
"grad_norm": 0.053466796875,
|
| 461 |
+
"learning_rate": 4.16e-05,
|
| 462 |
+
"loss": -0.0038484178483486177,
|
| 463 |
+
"num_tokens": 590152.0,
|
| 464 |
+
"reward": 0.13577499985694885,
|
| 465 |
+
"reward_std": 0.3495619535446167,
|
| 466 |
+
"rewards/dbre_reward/mean": 0.13577499985694885,
|
| 467 |
+
"rewards/dbre_reward/std": 0.34956197142601014,
|
| 468 |
+
"step": 85,
|
| 469 |
+
"step_time": 27.809819040997535
|
| 470 |
+
},
|
| 471 |
+
{
|
| 472 |
+
"clip_ratio/high_max": 0.0,
|
| 473 |
+
"clip_ratio/high_mean": 0.0,
|
| 474 |
+
"clip_ratio/low_mean": 0.0,
|
| 475 |
+
"clip_ratio/low_min": 0.0,
|
| 476 |
+
"clip_ratio/region_mean": 0.0,
|
| 477 |
+
"completions/clipped_ratio": 0.975,
|
| 478 |
+
"completions/max_length": 256.0,
|
| 479 |
+
"completions/max_terminated_length": 34.4,
|
| 480 |
+
"completions/mean_length": 253.7375,
|
| 481 |
+
"completions/mean_terminated_length": 33.1,
|
| 482 |
+
"completions/min_length": 236.6,
|
| 483 |
+
"completions/min_terminated_length": 31.8,
|
| 484 |
+
"entropy": 0.975049901008606,
|
| 485 |
+
"epoch": 1.8,
|
| 486 |
+
"frac_reward_zero_std": 0.3,
|
| 487 |
+
"grad_norm": 0.08447265625,
|
| 488 |
+
"learning_rate": 4.11e-05,
|
| 489 |
+
"loss": -0.0023170128464698792,
|
| 490 |
+
"num_tokens": 624931.0,
|
| 491 |
+
"reward": 0.14498749673366546,
|
| 492 |
+
"reward_std": 0.3355918139219284,
|
| 493 |
+
"rewards/dbre_reward/mean": 0.14498749673366546,
|
| 494 |
+
"rewards/dbre_reward/std": 0.3355918198823929,
|
| 495 |
+
"step": 90,
|
| 496 |
+
"step_time": 27.872812610585243
|
| 497 |
+
},
|
| 498 |
+
{
|
| 499 |
+
"clip_ratio/high_max": 0.0,
|
| 500 |
+
"clip_ratio/high_mean": 0.0,
|
| 501 |
+
"clip_ratio/low_mean": 0.0,
|
| 502 |
+
"clip_ratio/low_min": 0.0,
|
| 503 |
+
"clip_ratio/region_mean": 0.0,
|
| 504 |
+
"completions/clipped_ratio": 0.95,
|
| 505 |
+
"completions/max_length": 256.0,
|
| 506 |
+
"completions/max_terminated_length": 75.6,
|
| 507 |
+
"completions/mean_length": 250.775,
|
| 508 |
+
"completions/mean_terminated_length": 60.6,
|
| 509 |
+
"completions/min_length": 199.2,
|
| 510 |
+
"completions/min_terminated_length": 45.6,
|
| 511 |
+
"entropy": 0.9378853186964988,
|
| 512 |
+
"epoch": 1.9,
|
| 513 |
+
"frac_reward_zero_std": 0.4,
|
| 514 |
+
"grad_norm": 0.0576171875,
|
| 515 |
+
"learning_rate": 4.0600000000000004e-05,
|
| 516 |
+
"loss": -0.010608357191085816,
|
| 517 |
+
"num_tokens": 659473.0,
|
| 518 |
+
"reward": 0.13513749986886978,
|
| 519 |
+
"reward_std": 0.3332351267337799,
|
| 520 |
+
"rewards/dbre_reward/mean": 0.13513749986886978,
|
| 521 |
+
"rewards/dbre_reward/std": 0.3332351326942444,
|
| 522 |
+
"step": 95,
|
| 523 |
+
"step_time": 27.71230500699894
|
| 524 |
+
},
|
| 525 |
+
{
|
| 526 |
+
"clip_ratio/high_max": 0.0,
|
| 527 |
+
"clip_ratio/high_mean": 0.0,
|
| 528 |
+
"clip_ratio/low_mean": 0.0,
|
| 529 |
+
"clip_ratio/low_min": 0.0,
|
| 530 |
+
"clip_ratio/region_mean": 0.0,
|
| 531 |
+
"completions/clipped_ratio": 0.975,
|
| 532 |
+
"completions/max_length": 256.0,
|
| 533 |
+
"completions/max_terminated_length": 80.0,
|
| 534 |
+
"completions/mean_length": 254.6,
|
| 535 |
+
"completions/mean_terminated_length": 80.0,
|
| 536 |
+
"completions/min_length": 233.6,
|
| 537 |
+
"completions/min_terminated_length": 80.0,
|
| 538 |
+
"entropy": 0.8716862492263318,
|
| 539 |
+
"epoch": 2.0,
|
| 540 |
+
"frac_reward_zero_std": 0.4,
|
| 541 |
+
"grad_norm": 0.06396484375,
|
| 542 |
+
"learning_rate": 4.0100000000000006e-05,
|
| 543 |
+
"loss": 0.0028070926666259764,
|
| 544 |
+
"num_tokens": 694321.0,
|
| 545 |
+
"reward": 0.13687500059604646,
|
| 546 |
+
"reward_std": 0.33139119744300843,
|
| 547 |
+
"rewards/dbre_reward/mean": 0.13687500059604646,
|
| 548 |
+
"rewards/dbre_reward/std": 0.3313912093639374,
|
| 549 |
+
"step": 100,
|
| 550 |
+
"step_time": 27.905723336405934
|
| 551 |
+
},
|
| 552 |
+
{
|
| 553 |
+
"clip_ratio/high_max": 0.0,
|
| 554 |
+
"clip_ratio/high_mean": 0.0,
|
| 555 |
+
"clip_ratio/low_mean": 0.0,
|
| 556 |
+
"clip_ratio/low_min": 0.0,
|
| 557 |
+
"clip_ratio/region_mean": 0.0,
|
| 558 |
+
"completions/clipped_ratio": 0.95,
|
| 559 |
+
"completions/max_length": 256.0,
|
| 560 |
+
"completions/max_terminated_length": 138.6,
|
| 561 |
+
"completions/mean_length": 251.8625,
|
| 562 |
+
"completions/mean_terminated_length": 138.6,
|
| 563 |
+
"completions/min_length": 189.8,
|
| 564 |
+
"completions/min_terminated_length": 138.6,
|
| 565 |
+
"entropy": 0.9404824480414391,
|
| 566 |
+
"epoch": 2.1,
|
| 567 |
+
"frac_reward_zero_std": 0.5,
|
| 568 |
+
"grad_norm": 0.060791015625,
|
| 569 |
+
"learning_rate": 3.960000000000001e-05,
|
| 570 |
+
"loss": -0.004697377979755402,
|
| 571 |
+
"num_tokens": 728950.0,
|
| 572 |
+
"reward": 0.08519999980926514,
|
| 573 |
+
"reward_std": 0.24869290590286255,
|
| 574 |
+
"rewards/dbre_reward/mean": 0.08519999980926514,
|
| 575 |
+
"rewards/dbre_reward/std": 0.24869290590286255,
|
| 576 |
+
"step": 105,
|
| 577 |
+
"step_time": 27.779890004795742
|
| 578 |
+
},
|
| 579 |
+
{
|
| 580 |
+
"clip_ratio/high_max": 0.0,
|
| 581 |
+
"clip_ratio/high_mean": 0.0,
|
| 582 |
+
"clip_ratio/low_mean": 0.0,
|
| 583 |
+
"clip_ratio/low_min": 0.0,
|
| 584 |
+
"clip_ratio/region_mean": 0.0,
|
| 585 |
+
"completions/clipped_ratio": 0.9625,
|
| 586 |
+
"completions/max_length": 256.0,
|
| 587 |
+
"completions/max_terminated_length": 21.8,
|
| 588 |
+
"completions/mean_length": 248.425,
|
| 589 |
+
"completions/mean_terminated_length": 19.9,
|
| 590 |
+
"completions/min_length": 171.6,
|
| 591 |
+
"completions/min_terminated_length": 18.0,
|
| 592 |
+
"entropy": 1.0435175843536855,
|
| 593 |
+
"epoch": 2.2,
|
| 594 |
+
"frac_reward_zero_std": 0.3,
|
| 595 |
+
"grad_norm": 0.134765625,
|
| 596 |
+
"learning_rate": 3.91e-05,
|
| 597 |
+
"loss": -0.007500007748603821,
|
| 598 |
+
"num_tokens": 763304.0,
|
| 599 |
+
"reward": 0.13616250157356263,
|
| 600 |
+
"reward_std": 0.3243652701377869,
|
| 601 |
+
"rewards/dbre_reward/mean": 0.13616250157356263,
|
| 602 |
+
"rewards/dbre_reward/std": 0.3243652701377869,
|
| 603 |
+
"step": 110,
|
| 604 |
+
"step_time": 27.559503265400416
|
| 605 |
+
},
|
| 606 |
+
{
|
| 607 |
+
"clip_ratio/high_max": 0.0,
|
| 608 |
+
"clip_ratio/high_mean": 0.0,
|
| 609 |
+
"clip_ratio/low_mean": 0.0,
|
| 610 |
+
"clip_ratio/low_min": 0.0,
|
| 611 |
+
"clip_ratio/region_mean": 0.0,
|
| 612 |
+
"completions/clipped_ratio": 0.9875,
|
| 613 |
+
"completions/max_length": 256.0,
|
| 614 |
+
"completions/max_terminated_length": 10.0,
|
| 615 |
+
"completions/mean_length": 253.425,
|
| 616 |
+
"completions/mean_terminated_length": 10.0,
|
| 617 |
+
"completions/min_length": 214.8,
|
| 618 |
+
"completions/min_terminated_length": 10.0,
|
| 619 |
+
"entropy": 0.9914286866784096,
|
| 620 |
+
"epoch": 2.3,
|
| 621 |
+
"frac_reward_zero_std": 0.2,
|
| 622 |
+
"grad_norm": 0.0830078125,
|
| 623 |
+
"learning_rate": 3.86e-05,
|
| 624 |
+
"loss": -0.0037435129284858703,
|
| 625 |
+
"num_tokens": 798058.0,
|
| 626 |
+
"reward": 0.17202500104904175,
|
| 627 |
+
"reward_std": 0.37993268966674804,
|
| 628 |
+
"rewards/dbre_reward/mean": 0.17202500104904175,
|
| 629 |
+
"rewards/dbre_reward/std": 0.37993271350860597,
|
| 630 |
+
"step": 115,
|
| 631 |
+
"step_time": 27.579605548398103
|
| 632 |
+
},
|
| 633 |
+
{
|
| 634 |
+
"clip_ratio/high_max": 0.0,
|
| 635 |
+
"clip_ratio/high_mean": 0.0,
|
| 636 |
+
"clip_ratio/low_mean": 0.0,
|
| 637 |
+
"clip_ratio/low_min": 0.0,
|
| 638 |
+
"clip_ratio/region_mean": 0.0,
|
| 639 |
+
"completions/clipped_ratio": 0.95,
|
| 640 |
+
"completions/max_length": 256.0,
|
| 641 |
+
"completions/max_terminated_length": 128.6,
|
| 642 |
+
"completions/mean_length": 252.5,
|
| 643 |
+
"completions/mean_terminated_length": 117.6,
|
| 644 |
+
"completions/min_length": 209.0,
|
| 645 |
+
"completions/min_terminated_length": 106.6,
|
| 646 |
+
"entropy": 1.0246646717190742,
|
| 647 |
+
"epoch": 2.4,
|
| 648 |
+
"frac_reward_zero_std": 0.2,
|
| 649 |
+
"grad_norm": 0.0927734375,
|
| 650 |
+
"learning_rate": 3.8100000000000005e-05,
|
| 651 |
+
"loss": -0.004933140054345131,
|
| 652 |
+
"num_tokens": 832738.0,
|
| 653 |
+
"reward": 0.18355000019073486,
|
| 654 |
+
"reward_std": 0.36511489748954773,
|
| 655 |
+
"rewards/dbre_reward/mean": 0.18355000019073486,
|
| 656 |
+
"rewards/dbre_reward/std": 0.36511489748954773,
|
| 657 |
+
"step": 120,
|
| 658 |
+
"step_time": 27.58797151680337
|
| 659 |
+
},
|
| 660 |
+
{
|
| 661 |
+
"clip_ratio/high_max": 0.0,
|
| 662 |
+
"clip_ratio/high_mean": 0.0,
|
| 663 |
+
"clip_ratio/low_mean": 0.0,
|
| 664 |
+
"clip_ratio/low_min": 0.0,
|
| 665 |
+
"clip_ratio/region_mean": 0.0,
|
| 666 |
+
"completions/clipped_ratio": 0.9875,
|
| 667 |
+
"completions/max_length": 256.0,
|
| 668 |
+
"completions/max_terminated_length": 5.2,
|
| 669 |
+
"completions/mean_length": 253.125,
|
| 670 |
+
"completions/mean_terminated_length": 5.2,
|
| 671 |
+
"completions/min_length": 210.0,
|
| 672 |
+
"completions/min_terminated_length": 5.2,
|
| 673 |
+
"entropy": 0.9524292007088662,
|
| 674 |
+
"epoch": 2.5,
|
| 675 |
+
"frac_reward_zero_std": 0.3,
|
| 676 |
+
"grad_norm": 0.083984375,
|
| 677 |
+
"learning_rate": 3.76e-05,
|
| 678 |
+
"loss": -0.004205547273159027,
|
| 679 |
+
"num_tokens": 867468.0,
|
| 680 |
+
"reward": 0.22202499732375144,
|
| 681 |
+
"reward_std": 0.3661981552839279,
|
| 682 |
+
"rewards/dbre_reward/mean": 0.22202499732375144,
|
| 683 |
+
"rewards/dbre_reward/std": 0.3661981612443924,
|
| 684 |
+
"step": 125,
|
| 685 |
+
"step_time": 27.638399406394456
|
| 686 |
+
},
|
| 687 |
+
{
|
| 688 |
+
"clip_ratio/high_max": 0.0,
|
| 689 |
+
"clip_ratio/high_mean": 0.0,
|
| 690 |
+
"clip_ratio/low_mean": 0.0,
|
| 691 |
+
"clip_ratio/low_min": 0.0,
|
| 692 |
+
"clip_ratio/region_mean": 0.0,
|
| 693 |
+
"completions/clipped_ratio": 0.9875,
|
| 694 |
+
"completions/max_length": 256.0,
|
| 695 |
+
"completions/max_terminated_length": 34.4,
|
| 696 |
+
"completions/mean_length": 254.95,
|
| 697 |
+
"completions/mean_terminated_length": 34.4,
|
| 698 |
+
"completions/min_length": 239.2,
|
| 699 |
+
"completions/min_terminated_length": 34.4,
|
| 700 |
+
"entropy": 0.9268762037158013,
|
| 701 |
+
"epoch": 2.6,
|
| 702 |
+
"frac_reward_zero_std": 0.4,
|
| 703 |
+
"grad_norm": 0.060546875,
|
| 704 |
+
"learning_rate": 3.71e-05,
|
| 705 |
+
"loss": -0.002260996401309967,
|
| 706 |
+
"num_tokens": 902344.0,
|
| 707 |
+
"reward": 0.19577499628067016,
|
| 708 |
+
"reward_std": 0.3975414574146271,
|
| 709 |
+
"rewards/dbre_reward/mean": 0.19577499628067016,
|
| 710 |
+
"rewards/dbre_reward/std": 0.397541481256485,
|
| 711 |
+
"step": 130,
|
| 712 |
+
"step_time": 27.68072868139425
|
| 713 |
+
},
|
| 714 |
+
{
|
| 715 |
+
"clip_ratio/high_max": 0.0,
|
| 716 |
+
"clip_ratio/high_mean": 0.0,
|
| 717 |
+
"clip_ratio/low_mean": 0.0,
|
| 718 |
+
"clip_ratio/low_min": 0.0,
|
| 719 |
+
"clip_ratio/region_mean": 0.0,
|
| 720 |
+
"completions/clipped_ratio": 0.9875,
|
| 721 |
+
"completions/max_length": 256.0,
|
| 722 |
+
"completions/max_terminated_length": 26.6,
|
| 723 |
+
"completions/mean_length": 254.4625,
|
| 724 |
+
"completions/mean_terminated_length": 26.6,
|
| 725 |
+
"completions/min_length": 231.4,
|
| 726 |
+
"completions/min_terminated_length": 26.6,
|
| 727 |
+
"entropy": 0.9508431695401669,
|
| 728 |
+
"epoch": 2.7,
|
| 729 |
+
"frac_reward_zero_std": 0.1,
|
| 730 |
+
"grad_norm": 0.1025390625,
|
| 731 |
+
"learning_rate": 3.66e-05,
|
| 732 |
+
"loss": -0.004485464096069336,
|
| 733 |
+
"num_tokens": 937181.0,
|
| 734 |
+
"reward": 0.1941875010728836,
|
| 735 |
+
"reward_std": 0.39680722951889036,
|
| 736 |
+
"rewards/dbre_reward/mean": 0.1941875010728836,
|
| 737 |
+
"rewards/dbre_reward/std": 0.39680724740028384,
|
| 738 |
+
"step": 135,
|
| 739 |
+
"step_time": 27.79836461079831
|
| 740 |
+
},
|
| 741 |
+
{
|
| 742 |
+
"clip_ratio/high_max": 0.0,
|
| 743 |
+
"clip_ratio/high_mean": 0.0,
|
| 744 |
+
"clip_ratio/low_mean": 0.0,
|
| 745 |
+
"clip_ratio/low_min": 0.0,
|
| 746 |
+
"clip_ratio/region_mean": 0.0,
|
| 747 |
+
"completions/clipped_ratio": 1.0,
|
| 748 |
+
"completions/max_length": 256.0,
|
| 749 |
+
"completions/max_terminated_length": 0.0,
|
| 750 |
+
"completions/mean_length": 256.0,
|
| 751 |
+
"completions/mean_terminated_length": 0.0,
|
| 752 |
+
"completions/min_length": 256.0,
|
| 753 |
+
"completions/min_terminated_length": 0.0,
|
| 754 |
+
"entropy": 0.9166437476873398,
|
| 755 |
+
"epoch": 2.8,
|
| 756 |
+
"frac_reward_zero_std": 0.1,
|
| 757 |
+
"grad_norm": 0.078125,
|
| 758 |
+
"learning_rate": 3.61e-05,
|
| 759 |
+
"loss": 2.9802322387695314e-09,
|
| 760 |
+
"num_tokens": 972141.0,
|
| 761 |
+
"reward": 0.1720750018954277,
|
| 762 |
+
"reward_std": 0.3809880971908569,
|
| 763 |
+
"rewards/dbre_reward/mean": 0.1720750018954277,
|
| 764 |
+
"rewards/dbre_reward/std": 0.3809881091117859,
|
| 765 |
+
"step": 140,
|
| 766 |
+
"step_time": 27.89382164280105
|
| 767 |
+
},
|
| 768 |
+
{
|
| 769 |
+
"clip_ratio/high_max": 0.0,
|
| 770 |
+
"clip_ratio/high_mean": 0.0,
|
| 771 |
+
"clip_ratio/low_mean": 0.0,
|
| 772 |
+
"clip_ratio/low_min": 0.0,
|
| 773 |
+
"clip_ratio/region_mean": 0.0,
|
| 774 |
+
"completions/clipped_ratio": 0.95,
|
| 775 |
+
"completions/max_length": 256.0,
|
| 776 |
+
"completions/max_terminated_length": 117.2,
|
| 777 |
+
"completions/mean_length": 252.3125,
|
| 778 |
+
"completions/mean_terminated_length": 108.2,
|
| 779 |
+
"completions/min_length": 201.6,
|
| 780 |
+
"completions/min_terminated_length": 99.2,
|
| 781 |
+
"entropy": 1.039070624113083,
|
| 782 |
+
"epoch": 2.9,
|
| 783 |
+
"frac_reward_zero_std": 0.2,
|
| 784 |
+
"grad_norm": 0.08740234375,
|
| 785 |
+
"learning_rate": 3.56e-05,
|
| 786 |
+
"loss": 0.003438304364681244,
|
| 787 |
+
"num_tokens": 1006806.0,
|
| 788 |
+
"reward": 0.2208999961614609,
|
| 789 |
+
"reward_std": 0.401767635345459,
|
| 790 |
+
"rewards/dbre_reward/mean": 0.2208999961614609,
|
| 791 |
+
"rewards/dbre_reward/std": 0.4017676472663879,
|
| 792 |
+
"step": 145,
|
| 793 |
+
"step_time": 27.960635445202932
|
| 794 |
+
},
|
| 795 |
+
{
|
| 796 |
+
"clip_ratio/high_max": 0.0,
|
| 797 |
+
"clip_ratio/high_mean": 0.0,
|
| 798 |
+
"clip_ratio/low_mean": 0.0,
|
| 799 |
+
"clip_ratio/low_min": 0.0,
|
| 800 |
+
"clip_ratio/region_mean": 0.0,
|
| 801 |
+
"completions/clipped_ratio": 0.975,
|
| 802 |
+
"completions/max_length": 256.0,
|
| 803 |
+
"completions/max_terminated_length": 40.2,
|
| 804 |
+
"completions/mean_length": 253.0625,
|
| 805 |
+
"completions/mean_terminated_length": 27.7,
|
| 806 |
+
"completions/min_length": 220.0,
|
| 807 |
+
"completions/min_terminated_length": 15.2,
|
| 808 |
+
"entropy": 0.9523506201803684,
|
| 809 |
+
"epoch": 3.0,
|
| 810 |
+
"frac_reward_zero_std": 0.1,
|
| 811 |
+
"grad_norm": 0.0966796875,
|
| 812 |
+
"learning_rate": 3.51e-05,
|
| 813 |
+
"loss": -0.006567706167697906,
|
| 814 |
+
"num_tokens": 1041531.0,
|
| 815 |
+
"reward": 0.254237499833107,
|
| 816 |
+
"reward_std": 0.4319828271865845,
|
| 817 |
+
"rewards/dbre_reward/mean": 0.254237499833107,
|
| 818 |
+
"rewards/dbre_reward/std": 0.4319828271865845,
|
| 819 |
+
"step": 150,
|
| 820 |
+
"step_time": 33.83420682080032
|
| 821 |
+
},
|
| 822 |
+
{
|
| 823 |
+
"clip_ratio/high_max": 0.0,
|
| 824 |
+
"clip_ratio/high_mean": 0.0,
|
| 825 |
+
"clip_ratio/low_mean": 0.0,
|
| 826 |
+
"clip_ratio/low_min": 0.0,
|
| 827 |
+
"clip_ratio/region_mean": 0.0,
|
| 828 |
+
"completions/clipped_ratio": 0.95,
|
| 829 |
+
"completions/max_length": 256.0,
|
| 830 |
+
"completions/max_terminated_length": 135.8,
|
| 831 |
+
"completions/mean_length": 254.4875,
|
| 832 |
+
"completions/mean_terminated_length": 134.9,
|
| 833 |
+
"completions/min_length": 236.4,
|
| 834 |
+
"completions/min_terminated_length": 134.0,
|
| 835 |
+
"entropy": 0.9468083322048187,
|
| 836 |
+
"epoch": 3.1,
|
| 837 |
+
"frac_reward_zero_std": 0.2,
|
| 838 |
+
"grad_norm": 0.0703125,
|
| 839 |
+
"learning_rate": 3.46e-05,
|
| 840 |
+
"loss": -0.003655475750565529,
|
| 841 |
+
"num_tokens": 1076370.0,
|
| 842 |
+
"reward": 0.2434374988079071,
|
| 843 |
+
"reward_std": 0.4292252540588379,
|
| 844 |
+
"rewards/dbre_reward/mean": 0.2434374988079071,
|
| 845 |
+
"rewards/dbre_reward/std": 0.4292252600193024,
|
| 846 |
+
"step": 155,
|
| 847 |
+
"step_time": 35.7378796559904
|
| 848 |
+
},
|
| 849 |
+
{
|
| 850 |
+
"clip_ratio/high_max": 0.0,
|
| 851 |
+
"clip_ratio/high_mean": 0.0,
|
| 852 |
+
"clip_ratio/low_mean": 0.0,
|
| 853 |
+
"clip_ratio/low_min": 0.0,
|
| 854 |
+
"clip_ratio/region_mean": 0.0,
|
| 855 |
+
"completions/clipped_ratio": 0.9875,
|
| 856 |
+
"completions/max_length": 256.0,
|
| 857 |
+
"completions/max_terminated_length": 10.4,
|
| 858 |
+
"completions/mean_length": 253.45,
|
| 859 |
+
"completions/mean_terminated_length": 10.4,
|
| 860 |
+
"completions/min_length": 215.2,
|
| 861 |
+
"completions/min_terminated_length": 10.4,
|
| 862 |
+
"entropy": 0.9374438695609569,
|
| 863 |
+
"epoch": 3.2,
|
| 864 |
+
"frac_reward_zero_std": 0.3,
|
| 865 |
+
"grad_norm": 0.0693359375,
|
| 866 |
+
"learning_rate": 3.41e-05,
|
| 867 |
+
"loss": -0.005657447874546051,
|
| 868 |
+
"num_tokens": 1111126.0,
|
| 869 |
+
"reward": 0.2539374977350235,
|
| 870 |
+
"reward_std": 0.4268993496894836,
|
| 871 |
+
"rewards/dbre_reward/mean": 0.2539374977350235,
|
| 872 |
+
"rewards/dbre_reward/std": 0.4268993675708771,
|
| 873 |
+
"step": 160,
|
| 874 |
+
"step_time": 33.11068361419602
|
| 875 |
+
},
|
| 876 |
+
{
|
| 877 |
+
"clip_ratio/high_max": 0.0,
|
| 878 |
+
"clip_ratio/high_mean": 0.0,
|
| 879 |
+
"clip_ratio/low_mean": 0.0,
|
| 880 |
+
"clip_ratio/low_min": 0.0,
|
| 881 |
+
"clip_ratio/region_mean": 0.0,
|
| 882 |
+
"completions/clipped_ratio": 0.925,
|
| 883 |
+
"completions/max_length": 256.0,
|
| 884 |
+
"completions/max_terminated_length": 97.2,
|
| 885 |
+
"completions/mean_length": 244.7,
|
| 886 |
+
"completions/mean_terminated_length": 95.4,
|
| 887 |
+
"completions/min_length": 93.6,
|
| 888 |
+
"completions/min_terminated_length": 93.6,
|
| 889 |
+
"entropy": 1.0759656712412835,
|
| 890 |
+
"epoch": 3.3,
|
| 891 |
+
"frac_reward_zero_std": 0.1,
|
| 892 |
+
"grad_norm": 0.09716796875,
|
| 893 |
+
"learning_rate": 3.3600000000000004e-05,
|
| 894 |
+
"loss": 0.0013453811407089233,
|
| 895 |
+
"num_tokens": 1145182.0,
|
| 896 |
+
"reward": 0.1820499964058399,
|
| 897 |
+
"reward_std": 0.37800283133983614,
|
| 898 |
+
"rewards/dbre_reward/mean": 0.1820499964058399,
|
| 899 |
+
"rewards/dbre_reward/std": 0.37800286114215853,
|
| 900 |
+
"step": 165,
|
| 901 |
+
"step_time": 33.76342459159205
|
| 902 |
+
},
|
| 903 |
+
{
|
| 904 |
+
"clip_ratio/high_max": 0.0,
|
| 905 |
+
"clip_ratio/high_mean": 0.0,
|
| 906 |
+
"clip_ratio/low_mean": 0.0,
|
| 907 |
+
"clip_ratio/low_min": 0.0,
|
| 908 |
+
"clip_ratio/region_mean": 0.0,
|
| 909 |
+
"completions/clipped_ratio": 0.925,
|
| 910 |
+
"completions/max_length": 256.0,
|
| 911 |
+
"completions/max_terminated_length": 169.2,
|
| 912 |
+
"completions/mean_length": 250.3125,
|
| 913 |
+
"completions/mean_terminated_length": 168.7,
|
| 914 |
+
"completions/min_length": 168.2,
|
| 915 |
+
"completions/min_terminated_length": 168.2,
|
| 916 |
+
"entropy": 0.9453672260046005,
|
| 917 |
+
"epoch": 3.4,
|
| 918 |
+
"frac_reward_zero_std": 0.1,
|
| 919 |
+
"grad_norm": 0.0927734375,
|
| 920 |
+
"learning_rate": 3.3100000000000005e-05,
|
| 921 |
+
"loss": -0.001752069965004921,
|
| 922 |
+
"num_tokens": 1179687.0,
|
| 923 |
+
"reward": 0.21851249933242797,
|
| 924 |
+
"reward_std": 0.4150461137294769,
|
| 925 |
+
"rewards/dbre_reward/mean": 0.21851249933242797,
|
| 926 |
+
"rewards/dbre_reward/std": 0.4150461256504059,
|
| 927 |
+
"step": 170,
|
| 928 |
+
"step_time": 35.29977663640748
|
| 929 |
+
},
|
| 930 |
+
{
|
| 931 |
+
"clip_ratio/high_max": 0.0,
|
| 932 |
+
"clip_ratio/high_mean": 0.0,
|
| 933 |
+
"clip_ratio/low_mean": 0.0,
|
| 934 |
+
"clip_ratio/low_min": 0.0,
|
| 935 |
+
"clip_ratio/region_mean": 0.0,
|
| 936 |
+
"completions/clipped_ratio": 0.9625,
|
| 937 |
+
"completions/max_length": 256.0,
|
| 938 |
+
"completions/max_terminated_length": 82.6,
|
| 939 |
+
"completions/mean_length": 251.5625,
|
| 940 |
+
"completions/mean_terminated_length": 82.6,
|
| 941 |
+
"completions/min_length": 185.0,
|
| 942 |
+
"completions/min_terminated_length": 82.6,
|
| 943 |
+
"entropy": 0.9627169869840145,
|
| 944 |
+
"epoch": 3.5,
|
| 945 |
+
"frac_reward_zero_std": 0.0,
|
| 946 |
+
"grad_norm": 0.091796875,
|
| 947 |
+
"learning_rate": 3.26e-05,
|
| 948 |
+
"loss": -0.010714849084615707,
|
| 949 |
+
"num_tokens": 1214292.0,
|
| 950 |
+
"reward": 0.27426249384880064,
|
| 951 |
+
"reward_std": 0.42773920893669126,
|
| 952 |
+
"rewards/dbre_reward/mean": 0.27426249384880064,
|
| 953 |
+
"rewards/dbre_reward/std": 0.42773920893669126,
|
| 954 |
+
"step": 175,
|
| 955 |
+
"step_time": 32.60600872279319
|
| 956 |
+
},
|
| 957 |
+
{
|
| 958 |
+
"clip_ratio/high_max": 0.0,
|
| 959 |
+
"clip_ratio/high_mean": 0.0,
|
| 960 |
+
"clip_ratio/low_mean": 0.0,
|
| 961 |
+
"clip_ratio/low_min": 0.0,
|
| 962 |
+
"clip_ratio/region_mean": 0.0,
|
| 963 |
+
"completions/clipped_ratio": 0.9875,
|
| 964 |
+
"completions/max_length": 256.0,
|
| 965 |
+
"completions/max_terminated_length": 8.2,
|
| 966 |
+
"completions/mean_length": 253.3125,
|
| 967 |
+
"completions/mean_terminated_length": 8.2,
|
| 968 |
+
"completions/min_length": 213.0,
|
| 969 |
+
"completions/min_terminated_length": 8.2,
|
| 970 |
+
"entropy": 1.0290320612490178,
|
| 971 |
+
"epoch": 3.6,
|
| 972 |
+
"frac_reward_zero_std": 0.0,
|
| 973 |
+
"grad_norm": 0.09912109375,
|
| 974 |
+
"learning_rate": 3.21e-05,
|
| 975 |
+
"loss": -0.005978656560182571,
|
| 976 |
+
"num_tokens": 1249037.0,
|
| 977 |
+
"reward": 0.2384750008583069,
|
| 978 |
+
"reward_std": 0.4241094350814819,
|
| 979 |
+
"rewards/dbre_reward/mean": 0.2384750008583069,
|
| 980 |
+
"rewards/dbre_reward/std": 0.4241094350814819,
|
| 981 |
+
"step": 180,
|
| 982 |
+
"step_time": 33.03327611140848
|
| 983 |
+
},
|
| 984 |
+
{
|
| 985 |
+
"clip_ratio/high_max": 0.0,
|
| 986 |
+
"clip_ratio/high_mean": 0.0,
|
| 987 |
+
"clip_ratio/low_mean": 0.0,
|
| 988 |
+
"clip_ratio/low_min": 0.0,
|
| 989 |
+
"clip_ratio/region_mean": 0.0,
|
| 990 |
+
"completions/clipped_ratio": 0.925,
|
| 991 |
+
"completions/max_length": 256.0,
|
| 992 |
+
"completions/max_terminated_length": 132.8,
|
| 993 |
+
"completions/mean_length": 246.1875,
|
| 994 |
+
"completions/mean_terminated_length": 131.9,
|
| 995 |
+
"completions/min_length": 131.0,
|
| 996 |
+
"completions/min_terminated_length": 131.0,
|
| 997 |
+
"entropy": 0.92076805382967,
|
| 998 |
+
"epoch": 3.7,
|
| 999 |
+
"frac_reward_zero_std": 0.0,
|
| 1000 |
+
"grad_norm": 0.09619140625,
|
| 1001 |
+
"learning_rate": 3.16e-05,
|
| 1002 |
+
"loss": -0.009967343509197235,
|
| 1003 |
+
"num_tokens": 1283212.0,
|
| 1004 |
+
"reward": 0.33720000386238097,
|
| 1005 |
+
"reward_std": 0.46260204911231995,
|
| 1006 |
+
"rewards/dbre_reward/mean": 0.33720000386238097,
|
| 1007 |
+
"rewards/dbre_reward/std": 0.4626020550727844,
|
| 1008 |
+
"step": 185,
|
| 1009 |
+
"step_time": 34.82510126640555
|
| 1010 |
+
},
|
| 1011 |
+
{
|
| 1012 |
+
"clip_ratio/high_max": 0.0,
|
| 1013 |
+
"clip_ratio/high_mean": 0.0,
|
| 1014 |
+
"clip_ratio/low_mean": 0.0,
|
| 1015 |
+
"clip_ratio/low_min": 0.0,
|
| 1016 |
+
"clip_ratio/region_mean": 0.0,
|
| 1017 |
+
"completions/clipped_ratio": 0.95,
|
| 1018 |
+
"completions/max_length": 256.0,
|
| 1019 |
+
"completions/max_terminated_length": 74.2,
|
| 1020 |
+
"completions/mean_length": 248.875,
|
| 1021 |
+
"completions/mean_terminated_length": 45.4,
|
| 1022 |
+
"completions/min_length": 170.2,
|
| 1023 |
+
"completions/min_terminated_length": 16.6,
|
| 1024 |
+
"entropy": 1.0851470515131951,
|
| 1025 |
+
"epoch": 3.8,
|
| 1026 |
+
"frac_reward_zero_std": 0.2,
|
| 1027 |
+
"grad_norm": 0.09765625,
|
| 1028 |
+
"learning_rate": 3.1100000000000004e-05,
|
| 1029 |
+
"loss": 0.006586405634880066,
|
| 1030 |
+
"num_tokens": 1317602.0,
|
| 1031 |
+
"reward": 0.2771000027656555,
|
| 1032 |
+
"reward_std": 0.43828830122947693,
|
| 1033 |
+
"rewards/dbre_reward/mean": 0.2771000027656555,
|
| 1034 |
+
"rewards/dbre_reward/std": 0.43828831911087035,
|
| 1035 |
+
"step": 190,
|
| 1036 |
+
"step_time": 36.24654102979984
|
| 1037 |
+
},
|
| 1038 |
+
{
|
| 1039 |
+
"clip_ratio/high_max": 0.0,
|
| 1040 |
+
"clip_ratio/high_mean": 0.0,
|
| 1041 |
+
"clip_ratio/low_mean": 0.0,
|
| 1042 |
+
"clip_ratio/low_min": 0.0,
|
| 1043 |
+
"clip_ratio/region_mean": 0.0,
|
| 1044 |
+
"completions/clipped_ratio": 0.95,
|
| 1045 |
+
"completions/max_length": 256.0,
|
| 1046 |
+
"completions/max_terminated_length": 103.4,
|
| 1047 |
+
"completions/mean_length": 249.6625,
|
| 1048 |
+
"completions/mean_terminated_length": 103.4,
|
| 1049 |
+
"completions/min_length": 154.6,
|
| 1050 |
+
"completions/min_terminated_length": 103.4,
|
| 1051 |
+
"entropy": 0.9322984531521797,
|
| 1052 |
+
"epoch": 3.9,
|
| 1053 |
+
"frac_reward_zero_std": 0.0,
|
| 1054 |
+
"grad_norm": 0.099609375,
|
| 1055 |
+
"learning_rate": 3.06e-05,
|
| 1056 |
+
"loss": -0.011462598294019698,
|
| 1057 |
+
"num_tokens": 1352055.0,
|
| 1058 |
+
"reward": 0.30278749763965607,
|
| 1059 |
+
"reward_std": 0.4452593445777893,
|
| 1060 |
+
"rewards/dbre_reward/mean": 0.30278749763965607,
|
| 1061 |
+
"rewards/dbre_reward/std": 0.4452593445777893,
|
| 1062 |
+
"step": 195,
|
| 1063 |
+
"step_time": 29.027419628202914
|
| 1064 |
+
},
|
| 1065 |
+
{
|
| 1066 |
+
"clip_ratio/high_max": 0.0,
|
| 1067 |
+
"clip_ratio/high_mean": 0.0,
|
| 1068 |
+
"clip_ratio/low_mean": 0.0,
|
| 1069 |
+
"clip_ratio/low_min": 0.0,
|
| 1070 |
+
"clip_ratio/region_mean": 0.0,
|
| 1071 |
+
"completions/clipped_ratio": 0.975,
|
| 1072 |
+
"completions/max_length": 256.0,
|
| 1073 |
+
"completions/max_terminated_length": 21.8,
|
| 1074 |
+
"completions/mean_length": 250.9625,
|
| 1075 |
+
"completions/mean_terminated_length": 21.8,
|
| 1076 |
+
"completions/min_length": 175.4,
|
| 1077 |
+
"completions/min_terminated_length": 21.8,
|
| 1078 |
+
"entropy": 1.0117496035993099,
|
| 1079 |
+
"epoch": 4.0,
|
| 1080 |
+
"frac_reward_zero_std": 0.0,
|
| 1081 |
+
"grad_norm": 0.10009765625,
|
| 1082 |
+
"learning_rate": 3.01e-05,
|
| 1083 |
+
"loss": -0.017712239921092988,
|
| 1084 |
+
"num_tokens": 1386612.0,
|
| 1085 |
+
"reward": 0.2791749984025955,
|
| 1086 |
+
"reward_std": 0.43711588978767396,
|
| 1087 |
+
"rewards/dbre_reward/mean": 0.2791749984025955,
|
| 1088 |
+
"rewards/dbre_reward/std": 0.4371159017086029,
|
| 1089 |
+
"step": 200,
|
| 1090 |
+
"step_time": 27.529542883395333
|
| 1091 |
+
},
|
| 1092 |
+
{
|
| 1093 |
+
"clip_ratio/high_max": 0.0,
|
| 1094 |
+
"clip_ratio/high_mean": 0.0,
|
| 1095 |
+
"clip_ratio/low_mean": 0.0,
|
| 1096 |
+
"clip_ratio/low_min": 0.0,
|
| 1097 |
+
"clip_ratio/region_mean": 0.0,
|
| 1098 |
+
"completions/clipped_ratio": 0.9875,
|
| 1099 |
+
"completions/max_length": 256.0,
|
| 1100 |
+
"completions/max_terminated_length": 24.4,
|
| 1101 |
+
"completions/mean_length": 254.325,
|
| 1102 |
+
"completions/mean_terminated_length": 24.4,
|
| 1103 |
+
"completions/min_length": 229.2,
|
| 1104 |
+
"completions/min_terminated_length": 24.4,
|
| 1105 |
+
"entropy": 1.0013573169708252,
|
| 1106 |
+
"epoch": 4.1,
|
| 1107 |
+
"frac_reward_zero_std": 0.0,
|
| 1108 |
+
"grad_norm": 0.0908203125,
|
| 1109 |
+
"learning_rate": 2.96e-05,
|
| 1110 |
+
"loss": 0.006696997582912445,
|
| 1111 |
+
"num_tokens": 1421438.0,
|
| 1112 |
+
"reward": 0.333887505531311,
|
| 1113 |
+
"reward_std": 0.46158010959625245,
|
| 1114 |
+
"rewards/dbre_reward/mean": 0.333887505531311,
|
| 1115 |
+
"rewards/dbre_reward/std": 0.46158013343811033,
|
| 1116 |
+
"step": 205,
|
| 1117 |
+
"step_time": 27.625831649597966
|
| 1118 |
+
},
|
| 1119 |
+
{
|
| 1120 |
+
"clip_ratio/high_max": 0.0,
|
| 1121 |
+
"clip_ratio/high_mean": 0.0,
|
| 1122 |
+
"clip_ratio/low_mean": 0.0,
|
| 1123 |
+
"clip_ratio/low_min": 0.0,
|
| 1124 |
+
"clip_ratio/region_mean": 0.0,
|
| 1125 |
+
"completions/clipped_ratio": 1.0,
|
| 1126 |
+
"completions/max_length": 256.0,
|
| 1127 |
+
"completions/max_terminated_length": 0.0,
|
| 1128 |
+
"completions/mean_length": 256.0,
|
| 1129 |
+
"completions/mean_terminated_length": 0.0,
|
| 1130 |
+
"completions/min_length": 256.0,
|
| 1131 |
+
"completions/min_terminated_length": 0.0,
|
| 1132 |
+
"entropy": 1.0225385420024395,
|
| 1133 |
+
"epoch": 4.2,
|
| 1134 |
+
"frac_reward_zero_std": 0.0,
|
| 1135 |
+
"grad_norm": 0.1103515625,
|
| 1136 |
+
"learning_rate": 2.91e-05,
|
| 1137 |
+
"loss": -5.960464477539063e-09,
|
| 1138 |
+
"num_tokens": 1456398.0,
|
| 1139 |
+
"reward": 0.3235000044107437,
|
| 1140 |
+
"reward_std": 0.45962073802948,
|
| 1141 |
+
"rewards/dbre_reward/mean": 0.3235000044107437,
|
| 1142 |
+
"rewards/dbre_reward/std": 0.4596207320690155,
|
| 1143 |
+
"step": 210,
|
| 1144 |
+
"step_time": 27.678065907207202
|
| 1145 |
+
},
|
| 1146 |
+
{
|
| 1147 |
+
"clip_ratio/high_max": 0.0,
|
| 1148 |
+
"clip_ratio/high_mean": 0.0,
|
| 1149 |
+
"clip_ratio/low_mean": 0.0,
|
| 1150 |
+
"clip_ratio/low_min": 0.0,
|
| 1151 |
+
"clip_ratio/region_mean": 0.0,
|
| 1152 |
+
"completions/clipped_ratio": 0.9,
|
| 1153 |
+
"completions/max_length": 256.0,
|
| 1154 |
+
"completions/max_terminated_length": 153.2,
|
| 1155 |
+
"completions/mean_length": 244.6,
|
| 1156 |
+
"completions/mean_terminated_length": 113.6,
|
| 1157 |
+
"completions/min_length": 125.2,
|
| 1158 |
+
"completions/min_terminated_length": 74.0,
|
| 1159 |
+
"entropy": 1.0468589030206203,
|
| 1160 |
+
"epoch": 4.3,
|
| 1161 |
+
"frac_reward_zero_std": 0.3,
|
| 1162 |
+
"grad_norm": 0.07666015625,
|
| 1163 |
+
"learning_rate": 2.86e-05,
|
| 1164 |
+
"loss": -0.0048739627003669735,
|
| 1165 |
+
"num_tokens": 1490446.0,
|
| 1166 |
+
"reward": 0.22982499301433562,
|
| 1167 |
+
"reward_std": 0.4037831902503967,
|
| 1168 |
+
"rewards/dbre_reward/mean": 0.22982499301433562,
|
| 1169 |
+
"rewards/dbre_reward/std": 0.40378319621086123,
|
| 1170 |
+
"step": 215,
|
| 1171 |
+
"step_time": 27.696738472400465
|
| 1172 |
+
},
|
| 1173 |
+
{
|
| 1174 |
+
"clip_ratio/high_max": 0.0,
|
| 1175 |
+
"clip_ratio/high_mean": 0.0,
|
| 1176 |
+
"clip_ratio/low_mean": 0.0,
|
| 1177 |
+
"clip_ratio/low_min": 0.0,
|
| 1178 |
+
"clip_ratio/region_mean": 0.0,
|
| 1179 |
+
"completions/clipped_ratio": 0.9875,
|
| 1180 |
+
"completions/max_length": 256.0,
|
| 1181 |
+
"completions/max_terminated_length": 46.4,
|
| 1182 |
+
"completions/mean_length": 255.7,
|
| 1183 |
+
"completions/mean_terminated_length": 46.4,
|
| 1184 |
+
"completions/min_length": 251.2,
|
| 1185 |
+
"completions/min_terminated_length": 46.4,
|
| 1186 |
+
"entropy": 1.034738614410162,
|
| 1187 |
+
"epoch": 4.4,
|
| 1188 |
+
"frac_reward_zero_std": 0.0,
|
| 1189 |
+
"grad_norm": 0.10498046875,
|
| 1190 |
+
"learning_rate": 2.8100000000000005e-05,
|
| 1191 |
+
"loss": 0.0011565253138542176,
|
| 1192 |
+
"num_tokens": 1525382.0,
|
| 1193 |
+
"reward": 0.3366374969482422,
|
| 1194 |
+
"reward_std": 0.4732167422771454,
|
| 1195 |
+
"rewards/dbre_reward/mean": 0.3366374969482422,
|
| 1196 |
+
"rewards/dbre_reward/std": 0.47321674823760984,
|
| 1197 |
+
"step": 220,
|
| 1198 |
+
"step_time": 27.690395271993474
|
| 1199 |
+
},
|
| 1200 |
+
{
|
| 1201 |
+
"clip_ratio/high_max": 0.0,
|
| 1202 |
+
"clip_ratio/high_mean": 0.0,
|
| 1203 |
+
"clip_ratio/low_mean": 0.0,
|
| 1204 |
+
"clip_ratio/low_min": 0.0,
|
| 1205 |
+
"clip_ratio/region_mean": 0.0,
|
| 1206 |
+
"completions/clipped_ratio": 0.9875,
|
| 1207 |
+
"completions/max_length": 256.0,
|
| 1208 |
+
"completions/max_terminated_length": 27.0,
|
| 1209 |
+
"completions/mean_length": 254.4875,
|
| 1210 |
+
"completions/mean_terminated_length": 27.0,
|
| 1211 |
+
"completions/min_length": 231.8,
|
| 1212 |
+
"completions/min_terminated_length": 27.0,
|
| 1213 |
+
"entropy": 1.001887033134699,
|
| 1214 |
+
"epoch": 4.5,
|
| 1215 |
+
"frac_reward_zero_std": 0.1,
|
| 1216 |
+
"grad_norm": 0.1064453125,
|
| 1217 |
+
"learning_rate": 2.7600000000000003e-05,
|
| 1218 |
+
"loss": -0.004408703744411468,
|
| 1219 |
+
"num_tokens": 1560221.0,
|
| 1220 |
+
"reward": 0.3251750037074089,
|
| 1221 |
+
"reward_std": 0.4448351562023163,
|
| 1222 |
+
"rewards/dbre_reward/mean": 0.3251750037074089,
|
| 1223 |
+
"rewards/dbre_reward/std": 0.4448351800441742,
|
| 1224 |
+
"step": 225,
|
| 1225 |
+
"step_time": 27.640458112402122
|
| 1226 |
+
},
|
| 1227 |
+
{
|
| 1228 |
+
"clip_ratio/high_max": 0.0,
|
| 1229 |
+
"clip_ratio/high_mean": 0.0,
|
| 1230 |
+
"clip_ratio/low_mean": 0.0,
|
| 1231 |
+
"clip_ratio/low_min": 0.0,
|
| 1232 |
+
"clip_ratio/region_mean": 0.0,
|
| 1233 |
+
"completions/clipped_ratio": 0.9875,
|
| 1234 |
+
"completions/max_length": 256.0,
|
| 1235 |
+
"completions/max_terminated_length": 38.6,
|
| 1236 |
+
"completions/mean_length": 255.2125,
|
| 1237 |
+
"completions/mean_terminated_length": 38.6,
|
| 1238 |
+
"completions/min_length": 243.4,
|
| 1239 |
+
"completions/min_terminated_length": 38.6,
|
| 1240 |
+
"entropy": 0.9832980304956436,
|
| 1241 |
+
"epoch": 4.6,
|
| 1242 |
+
"frac_reward_zero_std": 0.1,
|
| 1243 |
+
"grad_norm": 0.09375,
|
| 1244 |
+
"learning_rate": 2.7100000000000005e-05,
|
| 1245 |
+
"loss": -0.0029202304780483247,
|
| 1246 |
+
"num_tokens": 1595118.0,
|
| 1247 |
+
"reward": 0.28887500166893004,
|
| 1248 |
+
"reward_std": 0.44461851716041567,
|
| 1249 |
+
"rewards/dbre_reward/mean": 0.28887500166893004,
|
| 1250 |
+
"rewards/dbre_reward/std": 0.4446185290813446,
|
| 1251 |
+
"step": 230,
|
| 1252 |
+
"step_time": 27.879659301796345
|
| 1253 |
+
},
|
| 1254 |
+
{
|
| 1255 |
+
"clip_ratio/high_max": 0.0,
|
| 1256 |
+
"clip_ratio/high_mean": 0.0,
|
| 1257 |
+
"clip_ratio/low_mean": 0.0,
|
| 1258 |
+
"clip_ratio/low_min": 0.0,
|
| 1259 |
+
"clip_ratio/region_mean": 0.0,
|
| 1260 |
+
"completions/clipped_ratio": 0.875,
|
| 1261 |
+
"completions/max_length": 256.0,
|
| 1262 |
+
"completions/max_terminated_length": 181.2,
|
| 1263 |
+
"completions/mean_length": 242.6625,
|
| 1264 |
+
"completions/mean_terminated_length": 157.15,
|
| 1265 |
+
"completions/min_length": 123.4,
|
| 1266 |
+
"completions/min_terminated_length": 123.4,
|
| 1267 |
+
"entropy": 1.0489521712064742,
|
| 1268 |
+
"epoch": 4.7,
|
| 1269 |
+
"frac_reward_zero_std": 0.1,
|
| 1270 |
+
"grad_norm": 0.10546875,
|
| 1271 |
+
"learning_rate": 2.6600000000000003e-05,
|
| 1272 |
+
"loss": -0.017240646481513976,
|
| 1273 |
+
"num_tokens": 1629011.0,
|
| 1274 |
+
"reward": 0.31321250200271605,
|
| 1275 |
+
"reward_std": 0.46374436616897585,
|
| 1276 |
+
"rewards/dbre_reward/mean": 0.31321250200271605,
|
| 1277 |
+
"rewards/dbre_reward/std": 0.46374437808990476,
|
| 1278 |
+
"step": 235,
|
| 1279 |
+
"step_time": 28.430774746800306
|
| 1280 |
+
},
|
| 1281 |
+
{
|
| 1282 |
+
"clip_ratio/high_max": 0.0,
|
| 1283 |
+
"clip_ratio/high_mean": 0.0,
|
| 1284 |
+
"clip_ratio/low_mean": 0.0,
|
| 1285 |
+
"clip_ratio/low_min": 0.0,
|
| 1286 |
+
"clip_ratio/region_mean": 0.0,
|
| 1287 |
+
"completions/clipped_ratio": 0.975,
|
| 1288 |
+
"completions/max_length": 256.0,
|
| 1289 |
+
"completions/max_terminated_length": 49.6,
|
| 1290 |
+
"completions/mean_length": 252.7,
|
| 1291 |
+
"completions/mean_terminated_length": 49.6,
|
| 1292 |
+
"completions/min_length": 203.2,
|
| 1293 |
+
"completions/min_terminated_length": 49.6,
|
| 1294 |
+
"entropy": 0.9966989070177078,
|
| 1295 |
+
"epoch": 4.8,
|
| 1296 |
+
"frac_reward_zero_std": 0.0,
|
| 1297 |
+
"grad_norm": 0.11083984375,
|
| 1298 |
+
"learning_rate": 2.61e-05,
|
| 1299 |
+
"loss": 0.0005952320992946625,
|
| 1300 |
+
"num_tokens": 1663707.0,
|
| 1301 |
+
"reward": 0.41918750405311583,
|
| 1302 |
+
"reward_std": 0.47992355227470396,
|
| 1303 |
+
"rewards/dbre_reward/mean": 0.41918750405311583,
|
| 1304 |
+
"rewards/dbre_reward/std": 0.47992355227470396,
|
| 1305 |
+
"step": 240,
|
| 1306 |
+
"step_time": 27.903805371007184
|
| 1307 |
+
},
|
| 1308 |
+
{
|
| 1309 |
+
"clip_ratio/high_max": 0.0,
|
| 1310 |
+
"clip_ratio/high_mean": 0.0,
|
| 1311 |
+
"clip_ratio/low_mean": 0.0,
|
| 1312 |
+
"clip_ratio/low_min": 0.0,
|
| 1313 |
+
"clip_ratio/region_mean": 0.0,
|
| 1314 |
+
"completions/clipped_ratio": 0.9625,
|
| 1315 |
+
"completions/max_length": 256.0,
|
| 1316 |
+
"completions/max_terminated_length": 43.8,
|
| 1317 |
+
"completions/mean_length": 250.95,
|
| 1318 |
+
"completions/mean_terminated_length": 38.3,
|
| 1319 |
+
"completions/min_length": 186.4,
|
| 1320 |
+
"completions/min_terminated_length": 32.8,
|
| 1321 |
+
"entropy": 0.9779959842562675,
|
| 1322 |
+
"epoch": 4.9,
|
| 1323 |
+
"frac_reward_zero_std": 0.0,
|
| 1324 |
+
"grad_norm": 0.083984375,
|
| 1325 |
+
"learning_rate": 2.5600000000000002e-05,
|
| 1326 |
+
"loss": -0.0041348889470100405,
|
| 1327 |
+
"num_tokens": 1698263.0,
|
| 1328 |
+
"reward": 0.3233749955892563,
|
| 1329 |
+
"reward_std": 0.4586354970932007,
|
| 1330 |
+
"rewards/dbre_reward/mean": 0.3233749955892563,
|
| 1331 |
+
"rewards/dbre_reward/std": 0.4586355030536652,
|
| 1332 |
+
"step": 245,
|
| 1333 |
+
"step_time": 27.984692039596847
|
| 1334 |
+
},
|
| 1335 |
+
{
|
| 1336 |
+
"clip_ratio/high_max": 0.0,
|
| 1337 |
+
"clip_ratio/high_mean": 0.0,
|
| 1338 |
+
"clip_ratio/low_mean": 0.0,
|
| 1339 |
+
"clip_ratio/low_min": 0.0,
|
| 1340 |
+
"clip_ratio/region_mean": 0.0,
|
| 1341 |
+
"completions/clipped_ratio": 0.9625,
|
| 1342 |
+
"completions/max_length": 256.0,
|
| 1343 |
+
"completions/max_terminated_length": 55.8,
|
| 1344 |
+
"completions/mean_length": 251.9125,
|
| 1345 |
+
"completions/mean_terminated_length": 47.0,
|
| 1346 |
+
"completions/min_length": 191.8,
|
| 1347 |
+
"completions/min_terminated_length": 38.2,
|
| 1348 |
+
"entropy": 0.9665410064160824,
|
| 1349 |
+
"epoch": 5.0,
|
| 1350 |
+
"frac_reward_zero_std": 0.1,
|
| 1351 |
+
"grad_norm": 0.10498046875,
|
| 1352 |
+
"learning_rate": 2.51e-05,
|
| 1353 |
+
"loss": 0.0026274655014276505,
|
| 1354 |
+
"num_tokens": 1732896.0,
|
| 1355 |
+
"reward": 0.37041249573230745,
|
| 1356 |
+
"reward_std": 0.46484237909317017,
|
| 1357 |
+
"rewards/dbre_reward/mean": 0.37041249573230745,
|
| 1358 |
+
"rewards/dbre_reward/std": 0.46484237909317017,
|
| 1359 |
+
"step": 250,
|
| 1360 |
+
"step_time": 27.929177726019407
|
| 1361 |
+
}
|
| 1362 |
+
],
|
| 1363 |
+
"logging_steps": 5,
|
| 1364 |
+
"max_steps": 500,
|
| 1365 |
+
"num_input_tokens_seen": 1732896,
|
| 1366 |
+
"num_train_epochs": 10,
|
| 1367 |
+
"save_steps": 50,
|
| 1368 |
+
"stateful_callbacks": {
|
| 1369 |
+
"TrainerControl": {
|
| 1370 |
+
"args": {
|
| 1371 |
+
"should_epoch_stop": false,
|
| 1372 |
+
"should_evaluate": false,
|
| 1373 |
+
"should_log": false,
|
| 1374 |
+
"should_save": true,
|
| 1375 |
+
"should_training_stop": false
|
| 1376 |
+
},
|
| 1377 |
+
"attributes": {}
|
| 1378 |
+
}
|
| 1379 |
+
},
|
| 1380 |
+
"total_flos": 0.0,
|
| 1381 |
+
"train_batch_size": 2,
|
| 1382 |
+
"trial_name": null,
|
| 1383 |
+
"trial_params": null
|
| 1384 |
+
}
|
dbre_trained/training_args.bin
ADDED
|
Binary file (7.12 kB). View file
|
|
|
elo_history/elo_history.json
CHANGED
|
The diff for this file is too large to render.
See raw diff
|
|
|
grpo_dbre/README.md
ADDED
|
@@ -0,0 +1,67 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
base_model: Qwen/Qwen2.5-Coder-1.5B-Instruct
|
| 3 |
+
library_name: transformers
|
| 4 |
+
model_name: grpo_dbre
|
| 5 |
+
tags:
|
| 6 |
+
- generated_from_trainer
|
| 7 |
+
- trl
|
| 8 |
+
- grpo
|
| 9 |
+
licence: license
|
| 10 |
+
---
|
| 11 |
+
|
| 12 |
+
# Model Card for grpo_dbre
|
| 13 |
+
|
| 14 |
+
This model is a fine-tuned version of [Qwen/Qwen2.5-Coder-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-Coder-1.5B-Instruct).
|
| 15 |
+
It has been trained using [TRL](https://github.com/huggingface/trl).
|
| 16 |
+
|
| 17 |
+
## Quick start
|
| 18 |
+
|
| 19 |
+
```python
|
| 20 |
+
from transformers import pipeline
|
| 21 |
+
|
| 22 |
+
question = "If you had a time machine, but could only go to the past or the future once and never return, which would you choose and why?"
|
| 23 |
+
generator = pipeline("text-generation", model="None", device="cuda")
|
| 24 |
+
output = generator([{"role": "user", "content": question}], max_new_tokens=128, return_full_text=False)[0]
|
| 25 |
+
print(output["generated_text"])
|
| 26 |
+
```
|
| 27 |
+
|
| 28 |
+
## Training procedure
|
| 29 |
+
|
| 30 |
+
|
| 31 |
+
|
| 32 |
+
|
| 33 |
+
|
| 34 |
+
This model was trained with GRPO, a method introduced in [DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models](https://huggingface.co/papers/2402.03300).
|
| 35 |
+
|
| 36 |
+
### Framework versions
|
| 37 |
+
|
| 38 |
+
- TRL: 1.2.0
|
| 39 |
+
- Transformers: 5.6.2
|
| 40 |
+
- Pytorch: 2.11.0
|
| 41 |
+
- Datasets: 4.8.4
|
| 42 |
+
- Tokenizers: 0.22.2
|
| 43 |
+
|
| 44 |
+
## Citations
|
| 45 |
+
|
| 46 |
+
Cite GRPO as:
|
| 47 |
+
|
| 48 |
+
```bibtex
|
| 49 |
+
@article{shao2024deepseekmath,
|
| 50 |
+
title = {{DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models}},
|
| 51 |
+
author = {Zhihong Shao and Peiyi Wang and Qihao Zhu and Runxin Xu and Junxiao Song and Mingchuan Zhang and Y. K. Li and Y. Wu and Daya Guo},
|
| 52 |
+
year = 2024,
|
| 53 |
+
eprint = {arXiv:2402.03300},
|
| 54 |
+
}
|
| 55 |
+
```
|
| 56 |
+
|
| 57 |
+
Cite TRL as:
|
| 58 |
+
|
| 59 |
+
```bibtex
|
| 60 |
+
@software{vonwerra2020trl,
|
| 61 |
+
title = {{TRL: Transformers Reinforcement Learning}},
|
| 62 |
+
author = {von Werra, Leandro and Belkada, Younes and Tunstall, Lewis and Beeching, Edward and Thrush, Tristan and Lambert, Nathan and Huang, Shengyi and Rasul, Kashif and Gallouédec, Quentin},
|
| 63 |
+
license = {Apache-2.0},
|
| 64 |
+
url = {https://github.com/huggingface/trl},
|
| 65 |
+
year = {2020}
|
| 66 |
+
}
|
| 67 |
+
```
|
grpo_dbre/checkpoint-100/README.md
ADDED
|
@@ -0,0 +1,209 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
base_model: Qwen/Qwen2.5-Coder-1.5B-Instruct
|
| 3 |
+
library_name: peft
|
| 4 |
+
pipeline_tag: text-generation
|
| 5 |
+
tags:
|
| 6 |
+
- base_model:adapter:Qwen/Qwen2.5-Coder-1.5B-Instruct
|
| 7 |
+
- grpo
|
| 8 |
+
- lora
|
| 9 |
+
- transformers
|
| 10 |
+
- trl
|
| 11 |
+
---
|
| 12 |
+
|
| 13 |
+
# Model Card for Model ID
|
| 14 |
+
|
| 15 |
+
<!-- Provide a quick summary of what the model is/does. -->
|
| 16 |
+
|
| 17 |
+
|
| 18 |
+
|
| 19 |
+
## Model Details
|
| 20 |
+
|
| 21 |
+
### Model Description
|
| 22 |
+
|
| 23 |
+
<!-- Provide a longer summary of what this model is. -->
|
| 24 |
+
|
| 25 |
+
|
| 26 |
+
|
| 27 |
+
- **Developed by:** [More Information Needed]
|
| 28 |
+
- **Funded by [optional]:** [More Information Needed]
|
| 29 |
+
- **Shared by [optional]:** [More Information Needed]
|
| 30 |
+
- **Model type:** [More Information Needed]
|
| 31 |
+
- **Language(s) (NLP):** [More Information Needed]
|
| 32 |
+
- **License:** [More Information Needed]
|
| 33 |
+
- **Finetuned from model [optional]:** [More Information Needed]
|
| 34 |
+
|
| 35 |
+
### Model Sources [optional]
|
| 36 |
+
|
| 37 |
+
<!-- Provide the basic links for the model. -->
|
| 38 |
+
|
| 39 |
+
- **Repository:** [More Information Needed]
|
| 40 |
+
- **Paper [optional]:** [More Information Needed]
|
| 41 |
+
- **Demo [optional]:** [More Information Needed]
|
| 42 |
+
|
| 43 |
+
## Uses
|
| 44 |
+
|
| 45 |
+
<!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
|
| 46 |
+
|
| 47 |
+
### Direct Use
|
| 48 |
+
|
| 49 |
+
<!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
|
| 50 |
+
|
| 51 |
+
[More Information Needed]
|
| 52 |
+
|
| 53 |
+
### Downstream Use [optional]
|
| 54 |
+
|
| 55 |
+
<!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
|
| 56 |
+
|
| 57 |
+
[More Information Needed]
|
| 58 |
+
|
| 59 |
+
### Out-of-Scope Use
|
| 60 |
+
|
| 61 |
+
<!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
|
| 62 |
+
|
| 63 |
+
[More Information Needed]
|
| 64 |
+
|
| 65 |
+
## Bias, Risks, and Limitations
|
| 66 |
+
|
| 67 |
+
<!-- This section is meant to convey both technical and sociotechnical limitations. -->
|
| 68 |
+
|
| 69 |
+
[More Information Needed]
|
| 70 |
+
|
| 71 |
+
### Recommendations
|
| 72 |
+
|
| 73 |
+
<!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
|
| 74 |
+
|
| 75 |
+
Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
|
| 76 |
+
|
| 77 |
+
## How to Get Started with the Model
|
| 78 |
+
|
| 79 |
+
Use the code below to get started with the model.
|
| 80 |
+
|
| 81 |
+
[More Information Needed]
|
| 82 |
+
|
| 83 |
+
## Training Details
|
| 84 |
+
|
| 85 |
+
### Training Data
|
| 86 |
+
|
| 87 |
+
<!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
|
| 88 |
+
|
| 89 |
+
[More Information Needed]
|
| 90 |
+
|
| 91 |
+
### Training Procedure
|
| 92 |
+
|
| 93 |
+
<!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
|
| 94 |
+
|
| 95 |
+
#### Preprocessing [optional]
|
| 96 |
+
|
| 97 |
+
[More Information Needed]
|
| 98 |
+
|
| 99 |
+
|
| 100 |
+
#### Training Hyperparameters
|
| 101 |
+
|
| 102 |
+
- **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
|
| 103 |
+
|
| 104 |
+
#### Speeds, Sizes, Times [optional]
|
| 105 |
+
|
| 106 |
+
<!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
|
| 107 |
+
|
| 108 |
+
[More Information Needed]
|
| 109 |
+
|
| 110 |
+
## Evaluation
|
| 111 |
+
|
| 112 |
+
<!-- This section describes the evaluation protocols and provides the results. -->
|
| 113 |
+
|
| 114 |
+
### Testing Data, Factors & Metrics
|
| 115 |
+
|
| 116 |
+
#### Testing Data
|
| 117 |
+
|
| 118 |
+
<!-- This should link to a Dataset Card if possible. -->
|
| 119 |
+
|
| 120 |
+
[More Information Needed]
|
| 121 |
+
|
| 122 |
+
#### Factors
|
| 123 |
+
|
| 124 |
+
<!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
|
| 125 |
+
|
| 126 |
+
[More Information Needed]
|
| 127 |
+
|
| 128 |
+
#### Metrics
|
| 129 |
+
|
| 130 |
+
<!-- These are the evaluation metrics being used, ideally with a description of why. -->
|
| 131 |
+
|
| 132 |
+
[More Information Needed]
|
| 133 |
+
|
| 134 |
+
### Results
|
| 135 |
+
|
| 136 |
+
[More Information Needed]
|
| 137 |
+
|
| 138 |
+
#### Summary
|
| 139 |
+
|
| 140 |
+
|
| 141 |
+
|
| 142 |
+
## Model Examination [optional]
|
| 143 |
+
|
| 144 |
+
<!-- Relevant interpretability work for the model goes here -->
|
| 145 |
+
|
| 146 |
+
[More Information Needed]
|
| 147 |
+
|
| 148 |
+
## Environmental Impact
|
| 149 |
+
|
| 150 |
+
<!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
|
| 151 |
+
|
| 152 |
+
Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
|
| 153 |
+
|
| 154 |
+
- **Hardware Type:** [More Information Needed]
|
| 155 |
+
- **Hours used:** [More Information Needed]
|
| 156 |
+
- **Cloud Provider:** [More Information Needed]
|
| 157 |
+
- **Compute Region:** [More Information Needed]
|
| 158 |
+
- **Carbon Emitted:** [More Information Needed]
|
| 159 |
+
|
| 160 |
+
## Technical Specifications [optional]
|
| 161 |
+
|
| 162 |
+
### Model Architecture and Objective
|
| 163 |
+
|
| 164 |
+
[More Information Needed]
|
| 165 |
+
|
| 166 |
+
### Compute Infrastructure
|
| 167 |
+
|
| 168 |
+
[More Information Needed]
|
| 169 |
+
|
| 170 |
+
#### Hardware
|
| 171 |
+
|
| 172 |
+
[More Information Needed]
|
| 173 |
+
|
| 174 |
+
#### Software
|
| 175 |
+
|
| 176 |
+
[More Information Needed]
|
| 177 |
+
|
| 178 |
+
## Citation [optional]
|
| 179 |
+
|
| 180 |
+
<!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
|
| 181 |
+
|
| 182 |
+
**BibTeX:**
|
| 183 |
+
|
| 184 |
+
[More Information Needed]
|
| 185 |
+
|
| 186 |
+
**APA:**
|
| 187 |
+
|
| 188 |
+
[More Information Needed]
|
| 189 |
+
|
| 190 |
+
## Glossary [optional]
|
| 191 |
+
|
| 192 |
+
<!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
|
| 193 |
+
|
| 194 |
+
[More Information Needed]
|
| 195 |
+
|
| 196 |
+
## More Information [optional]
|
| 197 |
+
|
| 198 |
+
[More Information Needed]
|
| 199 |
+
|
| 200 |
+
## Model Card Authors [optional]
|
| 201 |
+
|
| 202 |
+
[More Information Needed]
|
| 203 |
+
|
| 204 |
+
## Model Card Contact
|
| 205 |
+
|
| 206 |
+
[More Information Needed]
|
| 207 |
+
### Framework versions
|
| 208 |
+
|
| 209 |
+
- PEFT 0.19.1
|
grpo_dbre/checkpoint-100/adapter_config.json
ADDED
|
@@ -0,0 +1,45 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"alora_invocation_tokens": null,
|
| 3 |
+
"alpha_pattern": {},
|
| 4 |
+
"arrow_config": null,
|
| 5 |
+
"auto_mapping": null,
|
| 6 |
+
"base_model_name_or_path": "Qwen/Qwen2.5-Coder-1.5B-Instruct",
|
| 7 |
+
"bias": "none",
|
| 8 |
+
"corda_config": null,
|
| 9 |
+
"ensure_weight_tying": false,
|
| 10 |
+
"eva_config": null,
|
| 11 |
+
"exclude_modules": null,
|
| 12 |
+
"fan_in_fan_out": false,
|
| 13 |
+
"inference_mode": true,
|
| 14 |
+
"init_lora_weights": true,
|
| 15 |
+
"layer_replication": null,
|
| 16 |
+
"layers_pattern": null,
|
| 17 |
+
"layers_to_transform": null,
|
| 18 |
+
"loftq_config": {},
|
| 19 |
+
"lora_alpha": 16,
|
| 20 |
+
"lora_bias": false,
|
| 21 |
+
"lora_dropout": 0,
|
| 22 |
+
"lora_ga_config": null,
|
| 23 |
+
"megatron_config": null,
|
| 24 |
+
"megatron_core": "megatron.core",
|
| 25 |
+
"modules_to_save": null,
|
| 26 |
+
"peft_type": "LORA",
|
| 27 |
+
"peft_version": "0.19.1",
|
| 28 |
+
"qalora_group_size": 16,
|
| 29 |
+
"r": 16,
|
| 30 |
+
"rank_pattern": {},
|
| 31 |
+
"revision": null,
|
| 32 |
+
"target_modules": [
|
| 33 |
+
"q_proj",
|
| 34 |
+
"k_proj",
|
| 35 |
+
"o_proj",
|
| 36 |
+
"v_proj"
|
| 37 |
+
],
|
| 38 |
+
"target_parameters": null,
|
| 39 |
+
"task_type": "CAUSAL_LM",
|
| 40 |
+
"trainable_token_indices": null,
|
| 41 |
+
"use_bdlora": null,
|
| 42 |
+
"use_dora": false,
|
| 43 |
+
"use_qalora": false,
|
| 44 |
+
"use_rslora": false
|
| 45 |
+
}
|
grpo_dbre/checkpoint-100/chat_template.jinja
ADDED
|
@@ -0,0 +1,54 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{%- if tools %}
|
| 2 |
+
{{- '<|im_start|>system\n' }}
|
| 3 |
+
{%- if messages[0]['role'] == 'system' %}
|
| 4 |
+
{{- messages[0]['content'] }}
|
| 5 |
+
{%- else %}
|
| 6 |
+
{{- 'You are Qwen, created by Alibaba Cloud. You are a helpful assistant.' }}
|
| 7 |
+
{%- endif %}
|
| 8 |
+
{{- "\n\n# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>" }}
|
| 9 |
+
{%- for tool in tools %}
|
| 10 |
+
{{- "\n" }}
|
| 11 |
+
{{- tool | tojson }}
|
| 12 |
+
{%- endfor %}
|
| 13 |
+
{{- "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call><|im_end|>\n" }}
|
| 14 |
+
{%- else %}
|
| 15 |
+
{%- if messages[0]['role'] == 'system' %}
|
| 16 |
+
{{- '<|im_start|>system\n' + messages[0]['content'] + '<|im_end|>\n' }}
|
| 17 |
+
{%- else %}
|
| 18 |
+
{{- '<|im_start|>system\nYou are Qwen, created by Alibaba Cloud. You are a helpful assistant.<|im_end|>\n' }}
|
| 19 |
+
{%- endif %}
|
| 20 |
+
{%- endif %}
|
| 21 |
+
{%- for message in messages %}
|
| 22 |
+
{%- if (message.role == "user") or (message.role == "system" and not loop.first) or (message.role == "assistant" and not message.tool_calls) %}
|
| 23 |
+
{{- '<|im_start|>' + message.role + '\n' + message.content + '<|im_end|>' + '\n' }}
|
| 24 |
+
{%- elif message.role == "assistant" %}
|
| 25 |
+
{{- '<|im_start|>' + message.role }}
|
| 26 |
+
{%- if message.content %}
|
| 27 |
+
{{- '\n' + message.content }}
|
| 28 |
+
{%- endif %}
|
| 29 |
+
{%- for tool_call in message.tool_calls %}
|
| 30 |
+
{%- if tool_call.function is defined %}
|
| 31 |
+
{%- set tool_call = tool_call.function %}
|
| 32 |
+
{%- endif %}
|
| 33 |
+
{{- '\n<tool_call>\n{"name": "' }}
|
| 34 |
+
{{- tool_call.name }}
|
| 35 |
+
{{- '", "arguments": ' }}
|
| 36 |
+
{{- tool_call.arguments | tojson }}
|
| 37 |
+
{{- '}\n</tool_call>' }}
|
| 38 |
+
{%- endfor %}
|
| 39 |
+
{{- '<|im_end|>\n' }}
|
| 40 |
+
{%- elif message.role == "tool" %}
|
| 41 |
+
{%- if (loop.index0 == 0) or (messages[loop.index0 - 1].role != "tool") %}
|
| 42 |
+
{{- '<|im_start|>user' }}
|
| 43 |
+
{%- endif %}
|
| 44 |
+
{{- '\n<tool_response>\n' }}
|
| 45 |
+
{{- message.content }}
|
| 46 |
+
{{- '\n</tool_response>' }}
|
| 47 |
+
{%- if loop.last or (messages[loop.index0 + 1].role != "tool") %}
|
| 48 |
+
{{- '<|im_end|>\n' }}
|
| 49 |
+
{%- endif %}
|
| 50 |
+
{%- endif %}
|
| 51 |
+
{%- endfor %}
|
| 52 |
+
{%- if add_generation_prompt %}
|
| 53 |
+
{{- '<|im_start|>assistant\n' }}
|
| 54 |
+
{%- endif %}
|
grpo_dbre/checkpoint-100/rng_state.pth
ADDED
|
Binary file (14.6 kB). View file
|
|
|
grpo_dbre/checkpoint-100/scheduler.pt
ADDED
|
Binary file (1.47 kB). View file
|
|
|
grpo_dbre/checkpoint-100/tokenizer_config.json
ADDED
|
@@ -0,0 +1,30 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"add_prefix_space": false,
|
| 3 |
+
"backend": "tokenizers",
|
| 4 |
+
"bos_token": null,
|
| 5 |
+
"clean_up_tokenization_spaces": false,
|
| 6 |
+
"eos_token": "<|im_end|>",
|
| 7 |
+
"errors": "replace",
|
| 8 |
+
"extra_special_tokens": [
|
| 9 |
+
"<|im_start|>",
|
| 10 |
+
"<|im_end|>",
|
| 11 |
+
"<|object_ref_start|>",
|
| 12 |
+
"<|object_ref_end|>",
|
| 13 |
+
"<|box_start|>",
|
| 14 |
+
"<|box_end|>",
|
| 15 |
+
"<|quad_start|>",
|
| 16 |
+
"<|quad_end|>",
|
| 17 |
+
"<|vision_start|>",
|
| 18 |
+
"<|vision_end|>",
|
| 19 |
+
"<|vision_pad|>",
|
| 20 |
+
"<|image_pad|>",
|
| 21 |
+
"<|video_pad|>"
|
| 22 |
+
],
|
| 23 |
+
"is_local": false,
|
| 24 |
+
"local_files_only": false,
|
| 25 |
+
"model_max_length": 32768,
|
| 26 |
+
"pad_token": "<|im_end|>",
|
| 27 |
+
"split_special_tokens": false,
|
| 28 |
+
"tokenizer_class": "Qwen2Tokenizer",
|
| 29 |
+
"unk_token": null
|
| 30 |
+
}
|
grpo_dbre/checkpoint-100/trainer_state.json
ADDED
|
@@ -0,0 +1,574 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"best_global_step": null,
|
| 3 |
+
"best_metric": null,
|
| 4 |
+
"best_model_checkpoint": null,
|
| 5 |
+
"epoch": 2.0,
|
| 6 |
+
"eval_steps": 500,
|
| 7 |
+
"global_step": 100,
|
| 8 |
+
"is_hyper_param_search": false,
|
| 9 |
+
"is_local_process_zero": true,
|
| 10 |
+
"is_world_process_zero": true,
|
| 11 |
+
"log_history": [
|
| 12 |
+
{
|
| 13 |
+
"clip_ratio/high_max": 0.0,
|
| 14 |
+
"clip_ratio/high_mean": 0.0,
|
| 15 |
+
"clip_ratio/low_mean": 0.0,
|
| 16 |
+
"clip_ratio/low_min": 0.0,
|
| 17 |
+
"clip_ratio/region_mean": 0.0,
|
| 18 |
+
"completions/clipped_ratio": 0.9875,
|
| 19 |
+
"completions/max_length": 256.0,
|
| 20 |
+
"completions/max_terminated_length": 19.6,
|
| 21 |
+
"completions/mean_length": 254.025,
|
| 22 |
+
"completions/mean_terminated_length": 19.6,
|
| 23 |
+
"completions/min_length": 224.4,
|
| 24 |
+
"completions/min_terminated_length": 19.6,
|
| 25 |
+
"entropy": 0.903753462433815,
|
| 26 |
+
"epoch": 0.1,
|
| 27 |
+
"frac_reward_zero_std": 0.8,
|
| 28 |
+
"grad_norm": 0.046875,
|
| 29 |
+
"learning_rate": 4.96e-05,
|
| 30 |
+
"loss": -2.2351741790771484e-09,
|
| 31 |
+
"num_tokens": 34802.0,
|
| 32 |
+
"reward": 0.025,
|
| 33 |
+
"reward_std": 0.1,
|
| 34 |
+
"rewards/dbre_reward/mean": 0.025,
|
| 35 |
+
"rewards/dbre_reward/std": 0.1,
|
| 36 |
+
"step": 5,
|
| 37 |
+
"step_time": 27.396174477002933
|
| 38 |
+
},
|
| 39 |
+
{
|
| 40 |
+
"clip_ratio/high_max": 0.0,
|
| 41 |
+
"clip_ratio/high_mean": 0.0,
|
| 42 |
+
"clip_ratio/low_mean": 0.0,
|
| 43 |
+
"clip_ratio/low_min": 0.0,
|
| 44 |
+
"clip_ratio/region_mean": 0.0,
|
| 45 |
+
"completions/clipped_ratio": 0.975,
|
| 46 |
+
"completions/max_length": 256.0,
|
| 47 |
+
"completions/max_terminated_length": 48.4,
|
| 48 |
+
"completions/mean_length": 252.625,
|
| 49 |
+
"completions/mean_terminated_length": 48.4,
|
| 50 |
+
"completions/min_length": 202.0,
|
| 51 |
+
"completions/min_terminated_length": 48.4,
|
| 52 |
+
"entropy": 0.9422640666365624,
|
| 53 |
+
"epoch": 0.2,
|
| 54 |
+
"frac_reward_zero_std": 0.5,
|
| 55 |
+
"grad_norm": 0.0673828125,
|
| 56 |
+
"learning_rate": 4.91e-05,
|
| 57 |
+
"loss": -5.960464477539063e-09,
|
| 58 |
+
"num_tokens": 69492.0,
|
| 59 |
+
"reward": 0.09868749976158142,
|
| 60 |
+
"reward_std": 0.25848535895347596,
|
| 61 |
+
"rewards/dbre_reward/mean": 0.09868749976158142,
|
| 62 |
+
"rewards/dbre_reward/std": 0.2584853649139404,
|
| 63 |
+
"step": 10,
|
| 64 |
+
"step_time": 27.675077842207976
|
| 65 |
+
},
|
| 66 |
+
{
|
| 67 |
+
"clip_ratio/high_max": 0.0,
|
| 68 |
+
"clip_ratio/high_mean": 0.0,
|
| 69 |
+
"clip_ratio/low_mean": 0.0,
|
| 70 |
+
"clip_ratio/low_min": 0.0,
|
| 71 |
+
"clip_ratio/region_mean": 0.0,
|
| 72 |
+
"completions/clipped_ratio": 0.975,
|
| 73 |
+
"completions/max_length": 256.0,
|
| 74 |
+
"completions/max_terminated_length": 79.0,
|
| 75 |
+
"completions/mean_length": 254.5375,
|
| 76 |
+
"completions/mean_terminated_length": 79.0,
|
| 77 |
+
"completions/min_length": 232.6,
|
| 78 |
+
"completions/min_terminated_length": 79.0,
|
| 79 |
+
"entropy": 0.9289400212466716,
|
| 80 |
+
"epoch": 0.3,
|
| 81 |
+
"frac_reward_zero_std": 0.5,
|
| 82 |
+
"grad_norm": 0.059326171875,
|
| 83 |
+
"learning_rate": 4.86e-05,
|
| 84 |
+
"loss": -0.0020709306001663206,
|
| 85 |
+
"num_tokens": 104335.0,
|
| 86 |
+
"reward": 0.11038749814033508,
|
| 87 |
+
"reward_std": 0.27505697011947633,
|
| 88 |
+
"rewards/dbre_reward/mean": 0.11038749814033508,
|
| 89 |
+
"rewards/dbre_reward/std": 0.2750569820404053,
|
| 90 |
+
"step": 15,
|
| 91 |
+
"step_time": 27.678717108402633
|
| 92 |
+
},
|
| 93 |
+
{
|
| 94 |
+
"clip_ratio/high_max": 0.0,
|
| 95 |
+
"clip_ratio/high_mean": 0.0,
|
| 96 |
+
"clip_ratio/low_mean": 0.0,
|
| 97 |
+
"clip_ratio/low_min": 0.0,
|
| 98 |
+
"clip_ratio/region_mean": 0.0,
|
| 99 |
+
"completions/clipped_ratio": 0.9625,
|
| 100 |
+
"completions/max_length": 256.0,
|
| 101 |
+
"completions/max_terminated_length": 61.4,
|
| 102 |
+
"completions/mean_length": 250.2375,
|
| 103 |
+
"completions/mean_terminated_length": 61.4,
|
| 104 |
+
"completions/min_length": 163.8,
|
| 105 |
+
"completions/min_terminated_length": 61.4,
|
| 106 |
+
"entropy": 0.8890757068991662,
|
| 107 |
+
"epoch": 0.4,
|
| 108 |
+
"frac_reward_zero_std": 0.3,
|
| 109 |
+
"grad_norm": 0.052001953125,
|
| 110 |
+
"learning_rate": 4.8100000000000004e-05,
|
| 111 |
+
"loss": -0.00669153705239296,
|
| 112 |
+
"num_tokens": 138834.0,
|
| 113 |
+
"reward": 0.11190000027418137,
|
| 114 |
+
"reward_std": 0.3216355323791504,
|
| 115 |
+
"rewards/dbre_reward/mean": 0.11190000027418137,
|
| 116 |
+
"rewards/dbre_reward/std": 0.3216355502605438,
|
| 117 |
+
"step": 20,
|
| 118 |
+
"step_time": 27.725935825207852
|
| 119 |
+
},
|
| 120 |
+
{
|
| 121 |
+
"clip_ratio/high_max": 0.0,
|
| 122 |
+
"clip_ratio/high_mean": 0.0,
|
| 123 |
+
"clip_ratio/low_mean": 0.0,
|
| 124 |
+
"clip_ratio/low_min": 0.0,
|
| 125 |
+
"clip_ratio/region_mean": 0.0,
|
| 126 |
+
"completions/clipped_ratio": 0.975,
|
| 127 |
+
"completions/max_length": 256.0,
|
| 128 |
+
"completions/max_terminated_length": 78.2,
|
| 129 |
+
"completions/mean_length": 254.4875,
|
| 130 |
+
"completions/mean_terminated_length": 78.2,
|
| 131 |
+
"completions/min_length": 231.8,
|
| 132 |
+
"completions/min_terminated_length": 78.2,
|
| 133 |
+
"entropy": 0.906505486369133,
|
| 134 |
+
"epoch": 0.5,
|
| 135 |
+
"frac_reward_zero_std": 0.7,
|
| 136 |
+
"grad_norm": 0.0,
|
| 137 |
+
"learning_rate": 4.76e-05,
|
| 138 |
+
"loss": 0.00735630989074707,
|
| 139 |
+
"num_tokens": 173673.0,
|
| 140 |
+
"reward": 0.0375,
|
| 141 |
+
"reward_std": 0.15,
|
| 142 |
+
"rewards/dbre_reward/mean": 0.0375,
|
| 143 |
+
"rewards/dbre_reward/std": 0.15,
|
| 144 |
+
"step": 25,
|
| 145 |
+
"step_time": 27.84102148480888
|
| 146 |
+
},
|
| 147 |
+
{
|
| 148 |
+
"clip_ratio/high_max": 0.0,
|
| 149 |
+
"clip_ratio/high_mean": 0.0,
|
| 150 |
+
"clip_ratio/low_mean": 0.0,
|
| 151 |
+
"clip_ratio/low_min": 0.0,
|
| 152 |
+
"clip_ratio/region_mean": 0.0,
|
| 153 |
+
"completions/clipped_ratio": 0.9625,
|
| 154 |
+
"completions/max_length": 256.0,
|
| 155 |
+
"completions/max_terminated_length": 69.0,
|
| 156 |
+
"completions/mean_length": 251.2125,
|
| 157 |
+
"completions/mean_terminated_length": 56.1,
|
| 158 |
+
"completions/min_length": 196.8,
|
| 159 |
+
"completions/min_terminated_length": 43.2,
|
| 160 |
+
"entropy": 0.8586828224360943,
|
| 161 |
+
"epoch": 0.6,
|
| 162 |
+
"frac_reward_zero_std": 0.6,
|
| 163 |
+
"grad_norm": 0.059814453125,
|
| 164 |
+
"learning_rate": 4.71e-05,
|
| 165 |
+
"loss": -0.005647056177258492,
|
| 166 |
+
"num_tokens": 208250.0,
|
| 167 |
+
"reward": 0.0625,
|
| 168 |
+
"reward_std": 0.18662600517272948,
|
| 169 |
+
"rewards/dbre_reward/mean": 0.0625,
|
| 170 |
+
"rewards/dbre_reward/std": 0.18662601709365845,
|
| 171 |
+
"step": 30,
|
| 172 |
+
"step_time": 37.554382849001556
|
| 173 |
+
},
|
| 174 |
+
{
|
| 175 |
+
"clip_ratio/high_max": 0.0,
|
| 176 |
+
"clip_ratio/high_mean": 0.0,
|
| 177 |
+
"clip_ratio/low_mean": 0.0,
|
| 178 |
+
"clip_ratio/low_min": 0.0,
|
| 179 |
+
"clip_ratio/region_mean": 0.0,
|
| 180 |
+
"completions/clipped_ratio": 0.9625,
|
| 181 |
+
"completions/max_length": 256.0,
|
| 182 |
+
"completions/max_terminated_length": 115.6,
|
| 183 |
+
"completions/mean_length": 253.625,
|
| 184 |
+
"completions/mean_terminated_length": 115.6,
|
| 185 |
+
"completions/min_length": 218.0,
|
| 186 |
+
"completions/min_terminated_length": 115.6,
|
| 187 |
+
"entropy": 0.8725819021463395,
|
| 188 |
+
"epoch": 0.7,
|
| 189 |
+
"frac_reward_zero_std": 0.3,
|
| 190 |
+
"grad_norm": 0.06787109375,
|
| 191 |
+
"learning_rate": 4.660000000000001e-05,
|
| 192 |
+
"loss": 0.0011730872094631196,
|
| 193 |
+
"num_tokens": 243020.0,
|
| 194 |
+
"reward": 0.11033750027418136,
|
| 195 |
+
"reward_std": 0.31098498702049254,
|
| 196 |
+
"rewards/dbre_reward/mean": 0.11033750027418136,
|
| 197 |
+
"rewards/dbre_reward/std": 0.31098498702049254,
|
| 198 |
+
"step": 35,
|
| 199 |
+
"step_time": 35.878880111602484
|
| 200 |
+
},
|
| 201 |
+
{
|
| 202 |
+
"clip_ratio/high_max": 0.0,
|
| 203 |
+
"clip_ratio/high_mean": 0.0,
|
| 204 |
+
"clip_ratio/low_mean": 0.0,
|
| 205 |
+
"clip_ratio/low_min": 0.0,
|
| 206 |
+
"clip_ratio/region_mean": 0.0,
|
| 207 |
+
"completions/clipped_ratio": 1.0,
|
| 208 |
+
"completions/max_length": 256.0,
|
| 209 |
+
"completions/max_terminated_length": 0.0,
|
| 210 |
+
"completions/mean_length": 256.0,
|
| 211 |
+
"completions/mean_terminated_length": 0.0,
|
| 212 |
+
"completions/min_length": 256.0,
|
| 213 |
+
"completions/min_terminated_length": 0.0,
|
| 214 |
+
"entropy": 0.9090480573475361,
|
| 215 |
+
"epoch": 0.8,
|
| 216 |
+
"frac_reward_zero_std": 0.4,
|
| 217 |
+
"grad_norm": 0.046875,
|
| 218 |
+
"learning_rate": 4.61e-05,
|
| 219 |
+
"loss": -8.940696716308593e-09,
|
| 220 |
+
"num_tokens": 277980.0,
|
| 221 |
+
"reward": 0.09995000064373016,
|
| 222 |
+
"reward_std": 0.30480254292488096,
|
| 223 |
+
"rewards/dbre_reward/mean": 0.09995000064373016,
|
| 224 |
+
"rewards/dbre_reward/std": 0.3048025548458099,
|
| 225 |
+
"step": 40,
|
| 226 |
+
"step_time": 34.05516860120406
|
| 227 |
+
},
|
| 228 |
+
{
|
| 229 |
+
"clip_ratio/high_max": 0.0,
|
| 230 |
+
"clip_ratio/high_mean": 0.0,
|
| 231 |
+
"clip_ratio/low_mean": 0.0,
|
| 232 |
+
"clip_ratio/low_min": 0.0,
|
| 233 |
+
"clip_ratio/region_mean": 0.0,
|
| 234 |
+
"completions/clipped_ratio": 0.9875,
|
| 235 |
+
"completions/max_length": 256.0,
|
| 236 |
+
"completions/max_terminated_length": 26.0,
|
| 237 |
+
"completions/mean_length": 254.425,
|
| 238 |
+
"completions/mean_terminated_length": 26.0,
|
| 239 |
+
"completions/min_length": 230.8,
|
| 240 |
+
"completions/min_terminated_length": 26.0,
|
| 241 |
+
"entropy": 0.8977848328649998,
|
| 242 |
+
"epoch": 0.9,
|
| 243 |
+
"frac_reward_zero_std": 0.6,
|
| 244 |
+
"grad_norm": 0.0,
|
| 245 |
+
"learning_rate": 4.5600000000000004e-05,
|
| 246 |
+
"loss": 1.4901161193847657e-09,
|
| 247 |
+
"num_tokens": 312814.0,
|
| 248 |
+
"reward": 0.07415000200271607,
|
| 249 |
+
"reward_std": 0.197099506855011,
|
| 250 |
+
"rewards/dbre_reward/mean": 0.07415000200271607,
|
| 251 |
+
"rewards/dbre_reward/std": 0.197099506855011,
|
| 252 |
+
"step": 45,
|
| 253 |
+
"step_time": 34.819741847200206
|
| 254 |
+
},
|
| 255 |
+
{
|
| 256 |
+
"clip_ratio/high_max": 0.0,
|
| 257 |
+
"clip_ratio/high_mean": 0.0,
|
| 258 |
+
"clip_ratio/low_mean": 0.0,
|
| 259 |
+
"clip_ratio/low_min": 0.0,
|
| 260 |
+
"clip_ratio/region_mean": 0.0,
|
| 261 |
+
"completions/clipped_ratio": 0.975,
|
| 262 |
+
"completions/max_length": 256.0,
|
| 263 |
+
"completions/max_terminated_length": 20.4,
|
| 264 |
+
"completions/mean_length": 250.875,
|
| 265 |
+
"completions/mean_terminated_length": 20.4,
|
| 266 |
+
"completions/min_length": 174.0,
|
| 267 |
+
"completions/min_terminated_length": 20.4,
|
| 268 |
+
"entropy": 1.0103468239307403,
|
| 269 |
+
"epoch": 1.0,
|
| 270 |
+
"frac_reward_zero_std": 0.8,
|
| 271 |
+
"grad_norm": 0.0,
|
| 272 |
+
"learning_rate": 4.5100000000000005e-05,
|
| 273 |
+
"loss": -2.2351741790771484e-09,
|
| 274 |
+
"num_tokens": 347364.0,
|
| 275 |
+
"reward": 0.04707500040531158,
|
| 276 |
+
"reward_std": 0.12459058165550232,
|
| 277 |
+
"rewards/dbre_reward/mean": 0.04707500040531158,
|
| 278 |
+
"rewards/dbre_reward/std": 0.12459058165550232,
|
| 279 |
+
"step": 50,
|
| 280 |
+
"step_time": 36.14326306019211
|
| 281 |
+
},
|
| 282 |
+
{
|
| 283 |
+
"clip_ratio/high_max": 0.0,
|
| 284 |
+
"clip_ratio/high_mean": 0.0,
|
| 285 |
+
"clip_ratio/low_mean": 0.0,
|
| 286 |
+
"clip_ratio/low_min": 0.0,
|
| 287 |
+
"clip_ratio/region_mean": 0.0,
|
| 288 |
+
"completions/clipped_ratio": 0.9375,
|
| 289 |
+
"completions/max_length": 256.0,
|
| 290 |
+
"completions/max_terminated_length": 41.8,
|
| 291 |
+
"completions/mean_length": 244.5625,
|
| 292 |
+
"completions/mean_terminated_length": 18.55,
|
| 293 |
+
"completions/min_length": 161.4,
|
| 294 |
+
"completions/min_terminated_length": 7.8,
|
| 295 |
+
"entropy": 0.9930311724543571,
|
| 296 |
+
"epoch": 1.1,
|
| 297 |
+
"frac_reward_zero_std": 0.6,
|
| 298 |
+
"grad_norm": 0.06396484375,
|
| 299 |
+
"learning_rate": 4.46e-05,
|
| 300 |
+
"loss": -0.017940016090869905,
|
| 301 |
+
"num_tokens": 381409.0,
|
| 302 |
+
"reward": 0.06066250056028366,
|
| 303 |
+
"reward_std": 0.2124839812517166,
|
| 304 |
+
"rewards/dbre_reward/mean": 0.06066250056028366,
|
| 305 |
+
"rewards/dbre_reward/std": 0.21248398423194886,
|
| 306 |
+
"step": 55,
|
| 307 |
+
"step_time": 28.94829335878603
|
| 308 |
+
},
|
| 309 |
+
{
|
| 310 |
+
"clip_ratio/high_max": 0.0,
|
| 311 |
+
"clip_ratio/high_mean": 0.0,
|
| 312 |
+
"clip_ratio/low_mean": 0.0,
|
| 313 |
+
"clip_ratio/low_min": 0.0,
|
| 314 |
+
"clip_ratio/region_mean": 0.0,
|
| 315 |
+
"completions/clipped_ratio": 0.9875,
|
| 316 |
+
"completions/max_length": 256.0,
|
| 317 |
+
"completions/max_terminated_length": 9.6,
|
| 318 |
+
"completions/mean_length": 253.4,
|
| 319 |
+
"completions/mean_terminated_length": 9.6,
|
| 320 |
+
"completions/min_length": 214.4,
|
| 321 |
+
"completions/min_terminated_length": 9.6,
|
| 322 |
+
"entropy": 0.8150956228375434,
|
| 323 |
+
"epoch": 1.2,
|
| 324 |
+
"frac_reward_zero_std": 0.5,
|
| 325 |
+
"grad_norm": 0.06103515625,
|
| 326 |
+
"learning_rate": 4.41e-05,
|
| 327 |
+
"loss": -4.470348358154297e-09,
|
| 328 |
+
"num_tokens": 416161.0,
|
| 329 |
+
"reward": 0.11134999990463257,
|
| 330 |
+
"reward_std": 0.2740528523921967,
|
| 331 |
+
"rewards/dbre_reward/mean": 0.11134999990463257,
|
| 332 |
+
"rewards/dbre_reward/std": 0.27405285835266113,
|
| 333 |
+
"step": 60,
|
| 334 |
+
"step_time": 27.705245727201692
|
| 335 |
+
},
|
| 336 |
+
{
|
| 337 |
+
"clip_ratio/high_max": 0.0,
|
| 338 |
+
"clip_ratio/high_mean": 0.0,
|
| 339 |
+
"clip_ratio/low_mean": 0.0,
|
| 340 |
+
"clip_ratio/low_min": 0.0,
|
| 341 |
+
"clip_ratio/region_mean": 0.0,
|
| 342 |
+
"completions/clipped_ratio": 0.975,
|
| 343 |
+
"completions/max_length": 256.0,
|
| 344 |
+
"completions/max_terminated_length": 41.0,
|
| 345 |
+
"completions/mean_length": 252.1625,
|
| 346 |
+
"completions/mean_terminated_length": 41.0,
|
| 347 |
+
"completions/min_length": 194.6,
|
| 348 |
+
"completions/min_terminated_length": 41.0,
|
| 349 |
+
"entropy": 0.9309644259512424,
|
| 350 |
+
"epoch": 1.3,
|
| 351 |
+
"frac_reward_zero_std": 0.7,
|
| 352 |
+
"grad_norm": 0.0,
|
| 353 |
+
"learning_rate": 4.36e-05,
|
| 354 |
+
"loss": -8.940696716308593e-09,
|
| 355 |
+
"num_tokens": 450814.0,
|
| 356 |
+
"reward": 0.0875,
|
| 357 |
+
"reward_std": 0.21124515533447266,
|
| 358 |
+
"rewards/dbre_reward/mean": 0.0875,
|
| 359 |
+
"rewards/dbre_reward/std": 0.21124515533447266,
|
| 360 |
+
"step": 65,
|
| 361 |
+
"step_time": 27.76657635839365
|
| 362 |
+
},
|
| 363 |
+
{
|
| 364 |
+
"clip_ratio/high_max": 0.0,
|
| 365 |
+
"clip_ratio/high_mean": 0.0,
|
| 366 |
+
"clip_ratio/low_mean": 0.0,
|
| 367 |
+
"clip_ratio/low_min": 0.0,
|
| 368 |
+
"clip_ratio/region_mean": 0.0,
|
| 369 |
+
"completions/clipped_ratio": 0.9875,
|
| 370 |
+
"completions/max_length": 256.0,
|
| 371 |
+
"completions/max_terminated_length": 30.4,
|
| 372 |
+
"completions/mean_length": 254.7,
|
| 373 |
+
"completions/mean_terminated_length": 30.4,
|
| 374 |
+
"completions/min_length": 235.2,
|
| 375 |
+
"completions/min_terminated_length": 30.4,
|
| 376 |
+
"entropy": 0.8473479233682155,
|
| 377 |
+
"epoch": 1.4,
|
| 378 |
+
"frac_reward_zero_std": 0.6,
|
| 379 |
+
"grad_norm": 0.05078125,
|
| 380 |
+
"learning_rate": 4.3100000000000004e-05,
|
| 381 |
+
"loss": -2.2351741790771484e-09,
|
| 382 |
+
"num_tokens": 485670.0,
|
| 383 |
+
"reward": 0.07392499968409538,
|
| 384 |
+
"reward_std": 0.1946355789899826,
|
| 385 |
+
"rewards/dbre_reward/mean": 0.07392499968409538,
|
| 386 |
+
"rewards/dbre_reward/std": 0.1946355879306793,
|
| 387 |
+
"step": 70,
|
| 388 |
+
"step_time": 27.701584570185513
|
| 389 |
+
},
|
| 390 |
+
{
|
| 391 |
+
"clip_ratio/high_max": 0.0,
|
| 392 |
+
"clip_ratio/high_mean": 0.0,
|
| 393 |
+
"clip_ratio/low_mean": 0.0,
|
| 394 |
+
"clip_ratio/low_min": 0.0,
|
| 395 |
+
"clip_ratio/region_mean": 0.0,
|
| 396 |
+
"completions/clipped_ratio": 0.975,
|
| 397 |
+
"completions/max_length": 256.0,
|
| 398 |
+
"completions/max_terminated_length": 51.0,
|
| 399 |
+
"completions/mean_length": 252.7875,
|
| 400 |
+
"completions/mean_terminated_length": 51.0,
|
| 401 |
+
"completions/min_length": 204.6,
|
| 402 |
+
"completions/min_terminated_length": 51.0,
|
| 403 |
+
"entropy": 0.8822006396949291,
|
| 404 |
+
"epoch": 1.5,
|
| 405 |
+
"frac_reward_zero_std": 0.4,
|
| 406 |
+
"grad_norm": 0.046875,
|
| 407 |
+
"learning_rate": 4.26e-05,
|
| 408 |
+
"loss": -0.004587128758430481,
|
| 409 |
+
"num_tokens": 520373.0,
|
| 410 |
+
"reward": 0.09866249859333039,
|
| 411 |
+
"reward_std": 0.29479086995124815,
|
| 412 |
+
"rewards/dbre_reward/mean": 0.09866249859333039,
|
| 413 |
+
"rewards/dbre_reward/std": 0.29479087591171266,
|
| 414 |
+
"step": 75,
|
| 415 |
+
"step_time": 27.723117466596886
|
| 416 |
+
},
|
| 417 |
+
{
|
| 418 |
+
"clip_ratio/high_max": 0.0,
|
| 419 |
+
"clip_ratio/high_mean": 0.0,
|
| 420 |
+
"clip_ratio/low_mean": 0.0,
|
| 421 |
+
"clip_ratio/low_min": 0.0,
|
| 422 |
+
"clip_ratio/region_mean": 0.0,
|
| 423 |
+
"completions/clipped_ratio": 1.0,
|
| 424 |
+
"completions/max_length": 256.0,
|
| 425 |
+
"completions/max_terminated_length": 0.0,
|
| 426 |
+
"completions/mean_length": 256.0,
|
| 427 |
+
"completions/mean_terminated_length": 0.0,
|
| 428 |
+
"completions/min_length": 256.0,
|
| 429 |
+
"completions/min_terminated_length": 0.0,
|
| 430 |
+
"entropy": 0.8492010429501533,
|
| 431 |
+
"epoch": 1.6,
|
| 432 |
+
"frac_reward_zero_std": 0.4,
|
| 433 |
+
"grad_norm": 0.052734375,
|
| 434 |
+
"learning_rate": 4.21e-05,
|
| 435 |
+
"loss": -4.470348358154297e-09,
|
| 436 |
+
"num_tokens": 555333.0,
|
| 437 |
+
"reward": 0.13631249964237213,
|
| 438 |
+
"reward_std": 0.3453687012195587,
|
| 439 |
+
"rewards/dbre_reward/mean": 0.13631249964237213,
|
| 440 |
+
"rewards/dbre_reward/std": 0.34536872506141664,
|
| 441 |
+
"step": 80,
|
| 442 |
+
"step_time": 27.74355768300593
|
| 443 |
+
},
|
| 444 |
+
{
|
| 445 |
+
"clip_ratio/high_max": 0.0,
|
| 446 |
+
"clip_ratio/high_mean": 0.0,
|
| 447 |
+
"clip_ratio/low_mean": 0.0,
|
| 448 |
+
"clip_ratio/low_min": 0.0,
|
| 449 |
+
"clip_ratio/region_mean": 0.0,
|
| 450 |
+
"completions/clipped_ratio": 0.9875,
|
| 451 |
+
"completions/max_length": 256.0,
|
| 452 |
+
"completions/max_terminated_length": 23.0,
|
| 453 |
+
"completions/mean_length": 254.2375,
|
| 454 |
+
"completions/mean_terminated_length": 23.0,
|
| 455 |
+
"completions/min_length": 227.8,
|
| 456 |
+
"completions/min_terminated_length": 23.0,
|
| 457 |
+
"entropy": 0.9008926346898078,
|
| 458 |
+
"epoch": 1.7,
|
| 459 |
+
"frac_reward_zero_std": 0.4,
|
| 460 |
+
"grad_norm": 0.053466796875,
|
| 461 |
+
"learning_rate": 4.16e-05,
|
| 462 |
+
"loss": -0.0038484178483486177,
|
| 463 |
+
"num_tokens": 590152.0,
|
| 464 |
+
"reward": 0.13577499985694885,
|
| 465 |
+
"reward_std": 0.3495619535446167,
|
| 466 |
+
"rewards/dbre_reward/mean": 0.13577499985694885,
|
| 467 |
+
"rewards/dbre_reward/std": 0.34956197142601014,
|
| 468 |
+
"step": 85,
|
| 469 |
+
"step_time": 27.809819040997535
|
| 470 |
+
},
|
| 471 |
+
{
|
| 472 |
+
"clip_ratio/high_max": 0.0,
|
| 473 |
+
"clip_ratio/high_mean": 0.0,
|
| 474 |
+
"clip_ratio/low_mean": 0.0,
|
| 475 |
+
"clip_ratio/low_min": 0.0,
|
| 476 |
+
"clip_ratio/region_mean": 0.0,
|
| 477 |
+
"completions/clipped_ratio": 0.975,
|
| 478 |
+
"completions/max_length": 256.0,
|
| 479 |
+
"completions/max_terminated_length": 34.4,
|
| 480 |
+
"completions/mean_length": 253.7375,
|
| 481 |
+
"completions/mean_terminated_length": 33.1,
|
| 482 |
+
"completions/min_length": 236.6,
|
| 483 |
+
"completions/min_terminated_length": 31.8,
|
| 484 |
+
"entropy": 0.975049901008606,
|
| 485 |
+
"epoch": 1.8,
|
| 486 |
+
"frac_reward_zero_std": 0.3,
|
| 487 |
+
"grad_norm": 0.08447265625,
|
| 488 |
+
"learning_rate": 4.11e-05,
|
| 489 |
+
"loss": -0.0023170128464698792,
|
| 490 |
+
"num_tokens": 624931.0,
|
| 491 |
+
"reward": 0.14498749673366546,
|
| 492 |
+
"reward_std": 0.3355918139219284,
|
| 493 |
+
"rewards/dbre_reward/mean": 0.14498749673366546,
|
| 494 |
+
"rewards/dbre_reward/std": 0.3355918198823929,
|
| 495 |
+
"step": 90,
|
| 496 |
+
"step_time": 27.872812610585243
|
| 497 |
+
},
|
| 498 |
+
{
|
| 499 |
+
"clip_ratio/high_max": 0.0,
|
| 500 |
+
"clip_ratio/high_mean": 0.0,
|
| 501 |
+
"clip_ratio/low_mean": 0.0,
|
| 502 |
+
"clip_ratio/low_min": 0.0,
|
| 503 |
+
"clip_ratio/region_mean": 0.0,
|
| 504 |
+
"completions/clipped_ratio": 0.95,
|
| 505 |
+
"completions/max_length": 256.0,
|
| 506 |
+
"completions/max_terminated_length": 75.6,
|
| 507 |
+
"completions/mean_length": 250.775,
|
| 508 |
+
"completions/mean_terminated_length": 60.6,
|
| 509 |
+
"completions/min_length": 199.2,
|
| 510 |
+
"completions/min_terminated_length": 45.6,
|
| 511 |
+
"entropy": 0.9378853186964988,
|
| 512 |
+
"epoch": 1.9,
|
| 513 |
+
"frac_reward_zero_std": 0.4,
|
| 514 |
+
"grad_norm": 0.0576171875,
|
| 515 |
+
"learning_rate": 4.0600000000000004e-05,
|
| 516 |
+
"loss": -0.010608357191085816,
|
| 517 |
+
"num_tokens": 659473.0,
|
| 518 |
+
"reward": 0.13513749986886978,
|
| 519 |
+
"reward_std": 0.3332351267337799,
|
| 520 |
+
"rewards/dbre_reward/mean": 0.13513749986886978,
|
| 521 |
+
"rewards/dbre_reward/std": 0.3332351326942444,
|
| 522 |
+
"step": 95,
|
| 523 |
+
"step_time": 27.71230500699894
|
| 524 |
+
},
|
| 525 |
+
{
|
| 526 |
+
"clip_ratio/high_max": 0.0,
|
| 527 |
+
"clip_ratio/high_mean": 0.0,
|
| 528 |
+
"clip_ratio/low_mean": 0.0,
|
| 529 |
+
"clip_ratio/low_min": 0.0,
|
| 530 |
+
"clip_ratio/region_mean": 0.0,
|
| 531 |
+
"completions/clipped_ratio": 0.975,
|
| 532 |
+
"completions/max_length": 256.0,
|
| 533 |
+
"completions/max_terminated_length": 80.0,
|
| 534 |
+
"completions/mean_length": 254.6,
|
| 535 |
+
"completions/mean_terminated_length": 80.0,
|
| 536 |
+
"completions/min_length": 233.6,
|
| 537 |
+
"completions/min_terminated_length": 80.0,
|
| 538 |
+
"entropy": 0.8716862492263318,
|
| 539 |
+
"epoch": 2.0,
|
| 540 |
+
"frac_reward_zero_std": 0.4,
|
| 541 |
+
"grad_norm": 0.06396484375,
|
| 542 |
+
"learning_rate": 4.0100000000000006e-05,
|
| 543 |
+
"loss": 0.0028070926666259764,
|
| 544 |
+
"num_tokens": 694321.0,
|
| 545 |
+
"reward": 0.13687500059604646,
|
| 546 |
+
"reward_std": 0.33139119744300843,
|
| 547 |
+
"rewards/dbre_reward/mean": 0.13687500059604646,
|
| 548 |
+
"rewards/dbre_reward/std": 0.3313912093639374,
|
| 549 |
+
"step": 100,
|
| 550 |
+
"step_time": 27.905723336405934
|
| 551 |
+
}
|
| 552 |
+
],
|
| 553 |
+
"logging_steps": 5,
|
| 554 |
+
"max_steps": 500,
|
| 555 |
+
"num_input_tokens_seen": 694321,
|
| 556 |
+
"num_train_epochs": 10,
|
| 557 |
+
"save_steps": 50,
|
| 558 |
+
"stateful_callbacks": {
|
| 559 |
+
"TrainerControl": {
|
| 560 |
+
"args": {
|
| 561 |
+
"should_epoch_stop": false,
|
| 562 |
+
"should_evaluate": false,
|
| 563 |
+
"should_log": false,
|
| 564 |
+
"should_save": true,
|
| 565 |
+
"should_training_stop": false
|
| 566 |
+
},
|
| 567 |
+
"attributes": {}
|
| 568 |
+
}
|
| 569 |
+
},
|
| 570 |
+
"total_flos": 0.0,
|
| 571 |
+
"train_batch_size": 2,
|
| 572 |
+
"trial_name": null,
|
| 573 |
+
"trial_params": null
|
| 574 |
+
}
|
grpo_dbre/checkpoint-100/training_args.bin
ADDED
|
Binary file (7.12 kB). View file
|
|
|
grpo_dbre/checkpoint-150/README.md
ADDED
|
@@ -0,0 +1,209 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
base_model: Qwen/Qwen2.5-Coder-1.5B-Instruct
|
| 3 |
+
library_name: peft
|
| 4 |
+
pipeline_tag: text-generation
|
| 5 |
+
tags:
|
| 6 |
+
- base_model:adapter:Qwen/Qwen2.5-Coder-1.5B-Instruct
|
| 7 |
+
- grpo
|
| 8 |
+
- lora
|
| 9 |
+
- transformers
|
| 10 |
+
- trl
|
| 11 |
+
---
|
| 12 |
+
|
| 13 |
+
# Model Card for Model ID
|
| 14 |
+
|
| 15 |
+
<!-- Provide a quick summary of what the model is/does. -->
|
| 16 |
+
|
| 17 |
+
|
| 18 |
+
|
| 19 |
+
## Model Details
|
| 20 |
+
|
| 21 |
+
### Model Description
|
| 22 |
+
|
| 23 |
+
<!-- Provide a longer summary of what this model is. -->
|
| 24 |
+
|
| 25 |
+
|
| 26 |
+
|
| 27 |
+
- **Developed by:** [More Information Needed]
|
| 28 |
+
- **Funded by [optional]:** [More Information Needed]
|
| 29 |
+
- **Shared by [optional]:** [More Information Needed]
|
| 30 |
+
- **Model type:** [More Information Needed]
|
| 31 |
+
- **Language(s) (NLP):** [More Information Needed]
|
| 32 |
+
- **License:** [More Information Needed]
|
| 33 |
+
- **Finetuned from model [optional]:** [More Information Needed]
|
| 34 |
+
|
| 35 |
+
### Model Sources [optional]
|
| 36 |
+
|
| 37 |
+
<!-- Provide the basic links for the model. -->
|
| 38 |
+
|
| 39 |
+
- **Repository:** [More Information Needed]
|
| 40 |
+
- **Paper [optional]:** [More Information Needed]
|
| 41 |
+
- **Demo [optional]:** [More Information Needed]
|
| 42 |
+
|
| 43 |
+
## Uses
|
| 44 |
+
|
| 45 |
+
<!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
|
| 46 |
+
|
| 47 |
+
### Direct Use
|
| 48 |
+
|
| 49 |
+
<!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
|
| 50 |
+
|
| 51 |
+
[More Information Needed]
|
| 52 |
+
|
| 53 |
+
### Downstream Use [optional]
|
| 54 |
+
|
| 55 |
+
<!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
|
| 56 |
+
|
| 57 |
+
[More Information Needed]
|
| 58 |
+
|
| 59 |
+
### Out-of-Scope Use
|
| 60 |
+
|
| 61 |
+
<!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
|
| 62 |
+
|
| 63 |
+
[More Information Needed]
|
| 64 |
+
|
| 65 |
+
## Bias, Risks, and Limitations
|
| 66 |
+
|
| 67 |
+
<!-- This section is meant to convey both technical and sociotechnical limitations. -->
|
| 68 |
+
|
| 69 |
+
[More Information Needed]
|
| 70 |
+
|
| 71 |
+
### Recommendations
|
| 72 |
+
|
| 73 |
+
<!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
|
| 74 |
+
|
| 75 |
+
Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
|
| 76 |
+
|
| 77 |
+
## How to Get Started with the Model
|
| 78 |
+
|
| 79 |
+
Use the code below to get started with the model.
|
| 80 |
+
|
| 81 |
+
[More Information Needed]
|
| 82 |
+
|
| 83 |
+
## Training Details
|
| 84 |
+
|
| 85 |
+
### Training Data
|
| 86 |
+
|
| 87 |
+
<!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
|
| 88 |
+
|
| 89 |
+
[More Information Needed]
|
| 90 |
+
|
| 91 |
+
### Training Procedure
|
| 92 |
+
|
| 93 |
+
<!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
|
| 94 |
+
|
| 95 |
+
#### Preprocessing [optional]
|
| 96 |
+
|
| 97 |
+
[More Information Needed]
|
| 98 |
+
|
| 99 |
+
|
| 100 |
+
#### Training Hyperparameters
|
| 101 |
+
|
| 102 |
+
- **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
|
| 103 |
+
|
| 104 |
+
#### Speeds, Sizes, Times [optional]
|
| 105 |
+
|
| 106 |
+
<!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
|
| 107 |
+
|
| 108 |
+
[More Information Needed]
|
| 109 |
+
|
| 110 |
+
## Evaluation
|
| 111 |
+
|
| 112 |
+
<!-- This section describes the evaluation protocols and provides the results. -->
|
| 113 |
+
|
| 114 |
+
### Testing Data, Factors & Metrics
|
| 115 |
+
|
| 116 |
+
#### Testing Data
|
| 117 |
+
|
| 118 |
+
<!-- This should link to a Dataset Card if possible. -->
|
| 119 |
+
|
| 120 |
+
[More Information Needed]
|
| 121 |
+
|
| 122 |
+
#### Factors
|
| 123 |
+
|
| 124 |
+
<!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
|
| 125 |
+
|
| 126 |
+
[More Information Needed]
|
| 127 |
+
|
| 128 |
+
#### Metrics
|
| 129 |
+
|
| 130 |
+
<!-- These are the evaluation metrics being used, ideally with a description of why. -->
|
| 131 |
+
|
| 132 |
+
[More Information Needed]
|
| 133 |
+
|
| 134 |
+
### Results
|
| 135 |
+
|
| 136 |
+
[More Information Needed]
|
| 137 |
+
|
| 138 |
+
#### Summary
|
| 139 |
+
|
| 140 |
+
|
| 141 |
+
|
| 142 |
+
## Model Examination [optional]
|
| 143 |
+
|
| 144 |
+
<!-- Relevant interpretability work for the model goes here -->
|
| 145 |
+
|
| 146 |
+
[More Information Needed]
|
| 147 |
+
|
| 148 |
+
## Environmental Impact
|
| 149 |
+
|
| 150 |
+
<!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
|
| 151 |
+
|
| 152 |
+
Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
|
| 153 |
+
|
| 154 |
+
- **Hardware Type:** [More Information Needed]
|
| 155 |
+
- **Hours used:** [More Information Needed]
|
| 156 |
+
- **Cloud Provider:** [More Information Needed]
|
| 157 |
+
- **Compute Region:** [More Information Needed]
|
| 158 |
+
- **Carbon Emitted:** [More Information Needed]
|
| 159 |
+
|
| 160 |
+
## Technical Specifications [optional]
|
| 161 |
+
|
| 162 |
+
### Model Architecture and Objective
|
| 163 |
+
|
| 164 |
+
[More Information Needed]
|
| 165 |
+
|
| 166 |
+
### Compute Infrastructure
|
| 167 |
+
|
| 168 |
+
[More Information Needed]
|
| 169 |
+
|
| 170 |
+
#### Hardware
|
| 171 |
+
|
| 172 |
+
[More Information Needed]
|
| 173 |
+
|
| 174 |
+
#### Software
|
| 175 |
+
|
| 176 |
+
[More Information Needed]
|
| 177 |
+
|
| 178 |
+
## Citation [optional]
|
| 179 |
+
|
| 180 |
+
<!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
|
| 181 |
+
|
| 182 |
+
**BibTeX:**
|
| 183 |
+
|
| 184 |
+
[More Information Needed]
|
| 185 |
+
|
| 186 |
+
**APA:**
|
| 187 |
+
|
| 188 |
+
[More Information Needed]
|
| 189 |
+
|
| 190 |
+
## Glossary [optional]
|
| 191 |
+
|
| 192 |
+
<!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
|
| 193 |
+
|
| 194 |
+
[More Information Needed]
|
| 195 |
+
|
| 196 |
+
## More Information [optional]
|
| 197 |
+
|
| 198 |
+
[More Information Needed]
|
| 199 |
+
|
| 200 |
+
## Model Card Authors [optional]
|
| 201 |
+
|
| 202 |
+
[More Information Needed]
|
| 203 |
+
|
| 204 |
+
## Model Card Contact
|
| 205 |
+
|
| 206 |
+
[More Information Needed]
|
| 207 |
+
### Framework versions
|
| 208 |
+
|
| 209 |
+
- PEFT 0.19.1
|
grpo_dbre/checkpoint-150/adapter_config.json
ADDED
|
@@ -0,0 +1,45 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"alora_invocation_tokens": null,
|
| 3 |
+
"alpha_pattern": {},
|
| 4 |
+
"arrow_config": null,
|
| 5 |
+
"auto_mapping": null,
|
| 6 |
+
"base_model_name_or_path": "Qwen/Qwen2.5-Coder-1.5B-Instruct",
|
| 7 |
+
"bias": "none",
|
| 8 |
+
"corda_config": null,
|
| 9 |
+
"ensure_weight_tying": false,
|
| 10 |
+
"eva_config": null,
|
| 11 |
+
"exclude_modules": null,
|
| 12 |
+
"fan_in_fan_out": false,
|
| 13 |
+
"inference_mode": true,
|
| 14 |
+
"init_lora_weights": true,
|
| 15 |
+
"layer_replication": null,
|
| 16 |
+
"layers_pattern": null,
|
| 17 |
+
"layers_to_transform": null,
|
| 18 |
+
"loftq_config": {},
|
| 19 |
+
"lora_alpha": 16,
|
| 20 |
+
"lora_bias": false,
|
| 21 |
+
"lora_dropout": 0,
|
| 22 |
+
"lora_ga_config": null,
|
| 23 |
+
"megatron_config": null,
|
| 24 |
+
"megatron_core": "megatron.core",
|
| 25 |
+
"modules_to_save": null,
|
| 26 |
+
"peft_type": "LORA",
|
| 27 |
+
"peft_version": "0.19.1",
|
| 28 |
+
"qalora_group_size": 16,
|
| 29 |
+
"r": 16,
|
| 30 |
+
"rank_pattern": {},
|
| 31 |
+
"revision": null,
|
| 32 |
+
"target_modules": [
|
| 33 |
+
"q_proj",
|
| 34 |
+
"k_proj",
|
| 35 |
+
"o_proj",
|
| 36 |
+
"v_proj"
|
| 37 |
+
],
|
| 38 |
+
"target_parameters": null,
|
| 39 |
+
"task_type": "CAUSAL_LM",
|
| 40 |
+
"trainable_token_indices": null,
|
| 41 |
+
"use_bdlora": null,
|
| 42 |
+
"use_dora": false,
|
| 43 |
+
"use_qalora": false,
|
| 44 |
+
"use_rslora": false
|
| 45 |
+
}
|
grpo_dbre/checkpoint-150/chat_template.jinja
ADDED
|
@@ -0,0 +1,54 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{%- if tools %}
|
| 2 |
+
{{- '<|im_start|>system\n' }}
|
| 3 |
+
{%- if messages[0]['role'] == 'system' %}
|
| 4 |
+
{{- messages[0]['content'] }}
|
| 5 |
+
{%- else %}
|
| 6 |
+
{{- 'You are Qwen, created by Alibaba Cloud. You are a helpful assistant.' }}
|
| 7 |
+
{%- endif %}
|
| 8 |
+
{{- "\n\n# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>" }}
|
| 9 |
+
{%- for tool in tools %}
|
| 10 |
+
{{- "\n" }}
|
| 11 |
+
{{- tool | tojson }}
|
| 12 |
+
{%- endfor %}
|
| 13 |
+
{{- "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call><|im_end|>\n" }}
|
| 14 |
+
{%- else %}
|
| 15 |
+
{%- if messages[0]['role'] == 'system' %}
|
| 16 |
+
{{- '<|im_start|>system\n' + messages[0]['content'] + '<|im_end|>\n' }}
|
| 17 |
+
{%- else %}
|
| 18 |
+
{{- '<|im_start|>system\nYou are Qwen, created by Alibaba Cloud. You are a helpful assistant.<|im_end|>\n' }}
|
| 19 |
+
{%- endif %}
|
| 20 |
+
{%- endif %}
|
| 21 |
+
{%- for message in messages %}
|
| 22 |
+
{%- if (message.role == "user") or (message.role == "system" and not loop.first) or (message.role == "assistant" and not message.tool_calls) %}
|
| 23 |
+
{{- '<|im_start|>' + message.role + '\n' + message.content + '<|im_end|>' + '\n' }}
|
| 24 |
+
{%- elif message.role == "assistant" %}
|
| 25 |
+
{{- '<|im_start|>' + message.role }}
|
| 26 |
+
{%- if message.content %}
|
| 27 |
+
{{- '\n' + message.content }}
|
| 28 |
+
{%- endif %}
|
| 29 |
+
{%- for tool_call in message.tool_calls %}
|
| 30 |
+
{%- if tool_call.function is defined %}
|
| 31 |
+
{%- set tool_call = tool_call.function %}
|
| 32 |
+
{%- endif %}
|
| 33 |
+
{{- '\n<tool_call>\n{"name": "' }}
|
| 34 |
+
{{- tool_call.name }}
|
| 35 |
+
{{- '", "arguments": ' }}
|
| 36 |
+
{{- tool_call.arguments | tojson }}
|
| 37 |
+
{{- '}\n</tool_call>' }}
|
| 38 |
+
{%- endfor %}
|
| 39 |
+
{{- '<|im_end|>\n' }}
|
| 40 |
+
{%- elif message.role == "tool" %}
|
| 41 |
+
{%- if (loop.index0 == 0) or (messages[loop.index0 - 1].role != "tool") %}
|
| 42 |
+
{{- '<|im_start|>user' }}
|
| 43 |
+
{%- endif %}
|
| 44 |
+
{{- '\n<tool_response>\n' }}
|
| 45 |
+
{{- message.content }}
|
| 46 |
+
{{- '\n</tool_response>' }}
|
| 47 |
+
{%- if loop.last or (messages[loop.index0 + 1].role != "tool") %}
|
| 48 |
+
{{- '<|im_end|>\n' }}
|
| 49 |
+
{%- endif %}
|
| 50 |
+
{%- endif %}
|
| 51 |
+
{%- endfor %}
|
| 52 |
+
{%- if add_generation_prompt %}
|
| 53 |
+
{{- '<|im_start|>assistant\n' }}
|
| 54 |
+
{%- endif %}
|
grpo_dbre/checkpoint-150/rng_state.pth
ADDED
|
Binary file (14.6 kB). View file
|
|
|
grpo_dbre/checkpoint-150/scheduler.pt
ADDED
|
Binary file (1.47 kB). View file
|
|
|
grpo_dbre/checkpoint-150/tokenizer_config.json
ADDED
|
@@ -0,0 +1,30 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"add_prefix_space": false,
|
| 3 |
+
"backend": "tokenizers",
|
| 4 |
+
"bos_token": null,
|
| 5 |
+
"clean_up_tokenization_spaces": false,
|
| 6 |
+
"eos_token": "<|im_end|>",
|
| 7 |
+
"errors": "replace",
|
| 8 |
+
"extra_special_tokens": [
|
| 9 |
+
"<|im_start|>",
|
| 10 |
+
"<|im_end|>",
|
| 11 |
+
"<|object_ref_start|>",
|
| 12 |
+
"<|object_ref_end|>",
|
| 13 |
+
"<|box_start|>",
|
| 14 |
+
"<|box_end|>",
|
| 15 |
+
"<|quad_start|>",
|
| 16 |
+
"<|quad_end|>",
|
| 17 |
+
"<|vision_start|>",
|
| 18 |
+
"<|vision_end|>",
|
| 19 |
+
"<|vision_pad|>",
|
| 20 |
+
"<|image_pad|>",
|
| 21 |
+
"<|video_pad|>"
|
| 22 |
+
],
|
| 23 |
+
"is_local": false,
|
| 24 |
+
"local_files_only": false,
|
| 25 |
+
"model_max_length": 32768,
|
| 26 |
+
"pad_token": "<|im_end|>",
|
| 27 |
+
"split_special_tokens": false,
|
| 28 |
+
"tokenizer_class": "Qwen2Tokenizer",
|
| 29 |
+
"unk_token": null
|
| 30 |
+
}
|
grpo_dbre/checkpoint-150/trainer_state.json
ADDED
|
@@ -0,0 +1,844 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"best_global_step": null,
|
| 3 |
+
"best_metric": null,
|
| 4 |
+
"best_model_checkpoint": null,
|
| 5 |
+
"epoch": 3.0,
|
| 6 |
+
"eval_steps": 500,
|
| 7 |
+
"global_step": 150,
|
| 8 |
+
"is_hyper_param_search": false,
|
| 9 |
+
"is_local_process_zero": true,
|
| 10 |
+
"is_world_process_zero": true,
|
| 11 |
+
"log_history": [
|
| 12 |
+
{
|
| 13 |
+
"clip_ratio/high_max": 0.0,
|
| 14 |
+
"clip_ratio/high_mean": 0.0,
|
| 15 |
+
"clip_ratio/low_mean": 0.0,
|
| 16 |
+
"clip_ratio/low_min": 0.0,
|
| 17 |
+
"clip_ratio/region_mean": 0.0,
|
| 18 |
+
"completions/clipped_ratio": 0.9875,
|
| 19 |
+
"completions/max_length": 256.0,
|
| 20 |
+
"completions/max_terminated_length": 19.6,
|
| 21 |
+
"completions/mean_length": 254.025,
|
| 22 |
+
"completions/mean_terminated_length": 19.6,
|
| 23 |
+
"completions/min_length": 224.4,
|
| 24 |
+
"completions/min_terminated_length": 19.6,
|
| 25 |
+
"entropy": 0.903753462433815,
|
| 26 |
+
"epoch": 0.1,
|
| 27 |
+
"frac_reward_zero_std": 0.8,
|
| 28 |
+
"grad_norm": 0.046875,
|
| 29 |
+
"learning_rate": 4.96e-05,
|
| 30 |
+
"loss": -2.2351741790771484e-09,
|
| 31 |
+
"num_tokens": 34802.0,
|
| 32 |
+
"reward": 0.025,
|
| 33 |
+
"reward_std": 0.1,
|
| 34 |
+
"rewards/dbre_reward/mean": 0.025,
|
| 35 |
+
"rewards/dbre_reward/std": 0.1,
|
| 36 |
+
"step": 5,
|
| 37 |
+
"step_time": 27.396174477002933
|
| 38 |
+
},
|
| 39 |
+
{
|
| 40 |
+
"clip_ratio/high_max": 0.0,
|
| 41 |
+
"clip_ratio/high_mean": 0.0,
|
| 42 |
+
"clip_ratio/low_mean": 0.0,
|
| 43 |
+
"clip_ratio/low_min": 0.0,
|
| 44 |
+
"clip_ratio/region_mean": 0.0,
|
| 45 |
+
"completions/clipped_ratio": 0.975,
|
| 46 |
+
"completions/max_length": 256.0,
|
| 47 |
+
"completions/max_terminated_length": 48.4,
|
| 48 |
+
"completions/mean_length": 252.625,
|
| 49 |
+
"completions/mean_terminated_length": 48.4,
|
| 50 |
+
"completions/min_length": 202.0,
|
| 51 |
+
"completions/min_terminated_length": 48.4,
|
| 52 |
+
"entropy": 0.9422640666365624,
|
| 53 |
+
"epoch": 0.2,
|
| 54 |
+
"frac_reward_zero_std": 0.5,
|
| 55 |
+
"grad_norm": 0.0673828125,
|
| 56 |
+
"learning_rate": 4.91e-05,
|
| 57 |
+
"loss": -5.960464477539063e-09,
|
| 58 |
+
"num_tokens": 69492.0,
|
| 59 |
+
"reward": 0.09868749976158142,
|
| 60 |
+
"reward_std": 0.25848535895347596,
|
| 61 |
+
"rewards/dbre_reward/mean": 0.09868749976158142,
|
| 62 |
+
"rewards/dbre_reward/std": 0.2584853649139404,
|
| 63 |
+
"step": 10,
|
| 64 |
+
"step_time": 27.675077842207976
|
| 65 |
+
},
|
| 66 |
+
{
|
| 67 |
+
"clip_ratio/high_max": 0.0,
|
| 68 |
+
"clip_ratio/high_mean": 0.0,
|
| 69 |
+
"clip_ratio/low_mean": 0.0,
|
| 70 |
+
"clip_ratio/low_min": 0.0,
|
| 71 |
+
"clip_ratio/region_mean": 0.0,
|
| 72 |
+
"completions/clipped_ratio": 0.975,
|
| 73 |
+
"completions/max_length": 256.0,
|
| 74 |
+
"completions/max_terminated_length": 79.0,
|
| 75 |
+
"completions/mean_length": 254.5375,
|
| 76 |
+
"completions/mean_terminated_length": 79.0,
|
| 77 |
+
"completions/min_length": 232.6,
|
| 78 |
+
"completions/min_terminated_length": 79.0,
|
| 79 |
+
"entropy": 0.9289400212466716,
|
| 80 |
+
"epoch": 0.3,
|
| 81 |
+
"frac_reward_zero_std": 0.5,
|
| 82 |
+
"grad_norm": 0.059326171875,
|
| 83 |
+
"learning_rate": 4.86e-05,
|
| 84 |
+
"loss": -0.0020709306001663206,
|
| 85 |
+
"num_tokens": 104335.0,
|
| 86 |
+
"reward": 0.11038749814033508,
|
| 87 |
+
"reward_std": 0.27505697011947633,
|
| 88 |
+
"rewards/dbre_reward/mean": 0.11038749814033508,
|
| 89 |
+
"rewards/dbre_reward/std": 0.2750569820404053,
|
| 90 |
+
"step": 15,
|
| 91 |
+
"step_time": 27.678717108402633
|
| 92 |
+
},
|
| 93 |
+
{
|
| 94 |
+
"clip_ratio/high_max": 0.0,
|
| 95 |
+
"clip_ratio/high_mean": 0.0,
|
| 96 |
+
"clip_ratio/low_mean": 0.0,
|
| 97 |
+
"clip_ratio/low_min": 0.0,
|
| 98 |
+
"clip_ratio/region_mean": 0.0,
|
| 99 |
+
"completions/clipped_ratio": 0.9625,
|
| 100 |
+
"completions/max_length": 256.0,
|
| 101 |
+
"completions/max_terminated_length": 61.4,
|
| 102 |
+
"completions/mean_length": 250.2375,
|
| 103 |
+
"completions/mean_terminated_length": 61.4,
|
| 104 |
+
"completions/min_length": 163.8,
|
| 105 |
+
"completions/min_terminated_length": 61.4,
|
| 106 |
+
"entropy": 0.8890757068991662,
|
| 107 |
+
"epoch": 0.4,
|
| 108 |
+
"frac_reward_zero_std": 0.3,
|
| 109 |
+
"grad_norm": 0.052001953125,
|
| 110 |
+
"learning_rate": 4.8100000000000004e-05,
|
| 111 |
+
"loss": -0.00669153705239296,
|
| 112 |
+
"num_tokens": 138834.0,
|
| 113 |
+
"reward": 0.11190000027418137,
|
| 114 |
+
"reward_std": 0.3216355323791504,
|
| 115 |
+
"rewards/dbre_reward/mean": 0.11190000027418137,
|
| 116 |
+
"rewards/dbre_reward/std": 0.3216355502605438,
|
| 117 |
+
"step": 20,
|
| 118 |
+
"step_time": 27.725935825207852
|
| 119 |
+
},
|
| 120 |
+
{
|
| 121 |
+
"clip_ratio/high_max": 0.0,
|
| 122 |
+
"clip_ratio/high_mean": 0.0,
|
| 123 |
+
"clip_ratio/low_mean": 0.0,
|
| 124 |
+
"clip_ratio/low_min": 0.0,
|
| 125 |
+
"clip_ratio/region_mean": 0.0,
|
| 126 |
+
"completions/clipped_ratio": 0.975,
|
| 127 |
+
"completions/max_length": 256.0,
|
| 128 |
+
"completions/max_terminated_length": 78.2,
|
| 129 |
+
"completions/mean_length": 254.4875,
|
| 130 |
+
"completions/mean_terminated_length": 78.2,
|
| 131 |
+
"completions/min_length": 231.8,
|
| 132 |
+
"completions/min_terminated_length": 78.2,
|
| 133 |
+
"entropy": 0.906505486369133,
|
| 134 |
+
"epoch": 0.5,
|
| 135 |
+
"frac_reward_zero_std": 0.7,
|
| 136 |
+
"grad_norm": 0.0,
|
| 137 |
+
"learning_rate": 4.76e-05,
|
| 138 |
+
"loss": 0.00735630989074707,
|
| 139 |
+
"num_tokens": 173673.0,
|
| 140 |
+
"reward": 0.0375,
|
| 141 |
+
"reward_std": 0.15,
|
| 142 |
+
"rewards/dbre_reward/mean": 0.0375,
|
| 143 |
+
"rewards/dbre_reward/std": 0.15,
|
| 144 |
+
"step": 25,
|
| 145 |
+
"step_time": 27.84102148480888
|
| 146 |
+
},
|
| 147 |
+
{
|
| 148 |
+
"clip_ratio/high_max": 0.0,
|
| 149 |
+
"clip_ratio/high_mean": 0.0,
|
| 150 |
+
"clip_ratio/low_mean": 0.0,
|
| 151 |
+
"clip_ratio/low_min": 0.0,
|
| 152 |
+
"clip_ratio/region_mean": 0.0,
|
| 153 |
+
"completions/clipped_ratio": 0.9625,
|
| 154 |
+
"completions/max_length": 256.0,
|
| 155 |
+
"completions/max_terminated_length": 69.0,
|
| 156 |
+
"completions/mean_length": 251.2125,
|
| 157 |
+
"completions/mean_terminated_length": 56.1,
|
| 158 |
+
"completions/min_length": 196.8,
|
| 159 |
+
"completions/min_terminated_length": 43.2,
|
| 160 |
+
"entropy": 0.8586828224360943,
|
| 161 |
+
"epoch": 0.6,
|
| 162 |
+
"frac_reward_zero_std": 0.6,
|
| 163 |
+
"grad_norm": 0.059814453125,
|
| 164 |
+
"learning_rate": 4.71e-05,
|
| 165 |
+
"loss": -0.005647056177258492,
|
| 166 |
+
"num_tokens": 208250.0,
|
| 167 |
+
"reward": 0.0625,
|
| 168 |
+
"reward_std": 0.18662600517272948,
|
| 169 |
+
"rewards/dbre_reward/mean": 0.0625,
|
| 170 |
+
"rewards/dbre_reward/std": 0.18662601709365845,
|
| 171 |
+
"step": 30,
|
| 172 |
+
"step_time": 37.554382849001556
|
| 173 |
+
},
|
| 174 |
+
{
|
| 175 |
+
"clip_ratio/high_max": 0.0,
|
| 176 |
+
"clip_ratio/high_mean": 0.0,
|
| 177 |
+
"clip_ratio/low_mean": 0.0,
|
| 178 |
+
"clip_ratio/low_min": 0.0,
|
| 179 |
+
"clip_ratio/region_mean": 0.0,
|
| 180 |
+
"completions/clipped_ratio": 0.9625,
|
| 181 |
+
"completions/max_length": 256.0,
|
| 182 |
+
"completions/max_terminated_length": 115.6,
|
| 183 |
+
"completions/mean_length": 253.625,
|
| 184 |
+
"completions/mean_terminated_length": 115.6,
|
| 185 |
+
"completions/min_length": 218.0,
|
| 186 |
+
"completions/min_terminated_length": 115.6,
|
| 187 |
+
"entropy": 0.8725819021463395,
|
| 188 |
+
"epoch": 0.7,
|
| 189 |
+
"frac_reward_zero_std": 0.3,
|
| 190 |
+
"grad_norm": 0.06787109375,
|
| 191 |
+
"learning_rate": 4.660000000000001e-05,
|
| 192 |
+
"loss": 0.0011730872094631196,
|
| 193 |
+
"num_tokens": 243020.0,
|
| 194 |
+
"reward": 0.11033750027418136,
|
| 195 |
+
"reward_std": 0.31098498702049254,
|
| 196 |
+
"rewards/dbre_reward/mean": 0.11033750027418136,
|
| 197 |
+
"rewards/dbre_reward/std": 0.31098498702049254,
|
| 198 |
+
"step": 35,
|
| 199 |
+
"step_time": 35.878880111602484
|
| 200 |
+
},
|
| 201 |
+
{
|
| 202 |
+
"clip_ratio/high_max": 0.0,
|
| 203 |
+
"clip_ratio/high_mean": 0.0,
|
| 204 |
+
"clip_ratio/low_mean": 0.0,
|
| 205 |
+
"clip_ratio/low_min": 0.0,
|
| 206 |
+
"clip_ratio/region_mean": 0.0,
|
| 207 |
+
"completions/clipped_ratio": 1.0,
|
| 208 |
+
"completions/max_length": 256.0,
|
| 209 |
+
"completions/max_terminated_length": 0.0,
|
| 210 |
+
"completions/mean_length": 256.0,
|
| 211 |
+
"completions/mean_terminated_length": 0.0,
|
| 212 |
+
"completions/min_length": 256.0,
|
| 213 |
+
"completions/min_terminated_length": 0.0,
|
| 214 |
+
"entropy": 0.9090480573475361,
|
| 215 |
+
"epoch": 0.8,
|
| 216 |
+
"frac_reward_zero_std": 0.4,
|
| 217 |
+
"grad_norm": 0.046875,
|
| 218 |
+
"learning_rate": 4.61e-05,
|
| 219 |
+
"loss": -8.940696716308593e-09,
|
| 220 |
+
"num_tokens": 277980.0,
|
| 221 |
+
"reward": 0.09995000064373016,
|
| 222 |
+
"reward_std": 0.30480254292488096,
|
| 223 |
+
"rewards/dbre_reward/mean": 0.09995000064373016,
|
| 224 |
+
"rewards/dbre_reward/std": 0.3048025548458099,
|
| 225 |
+
"step": 40,
|
| 226 |
+
"step_time": 34.05516860120406
|
| 227 |
+
},
|
| 228 |
+
{
|
| 229 |
+
"clip_ratio/high_max": 0.0,
|
| 230 |
+
"clip_ratio/high_mean": 0.0,
|
| 231 |
+
"clip_ratio/low_mean": 0.0,
|
| 232 |
+
"clip_ratio/low_min": 0.0,
|
| 233 |
+
"clip_ratio/region_mean": 0.0,
|
| 234 |
+
"completions/clipped_ratio": 0.9875,
|
| 235 |
+
"completions/max_length": 256.0,
|
| 236 |
+
"completions/max_terminated_length": 26.0,
|
| 237 |
+
"completions/mean_length": 254.425,
|
| 238 |
+
"completions/mean_terminated_length": 26.0,
|
| 239 |
+
"completions/min_length": 230.8,
|
| 240 |
+
"completions/min_terminated_length": 26.0,
|
| 241 |
+
"entropy": 0.8977848328649998,
|
| 242 |
+
"epoch": 0.9,
|
| 243 |
+
"frac_reward_zero_std": 0.6,
|
| 244 |
+
"grad_norm": 0.0,
|
| 245 |
+
"learning_rate": 4.5600000000000004e-05,
|
| 246 |
+
"loss": 1.4901161193847657e-09,
|
| 247 |
+
"num_tokens": 312814.0,
|
| 248 |
+
"reward": 0.07415000200271607,
|
| 249 |
+
"reward_std": 0.197099506855011,
|
| 250 |
+
"rewards/dbre_reward/mean": 0.07415000200271607,
|
| 251 |
+
"rewards/dbre_reward/std": 0.197099506855011,
|
| 252 |
+
"step": 45,
|
| 253 |
+
"step_time": 34.819741847200206
|
| 254 |
+
},
|
| 255 |
+
{
|
| 256 |
+
"clip_ratio/high_max": 0.0,
|
| 257 |
+
"clip_ratio/high_mean": 0.0,
|
| 258 |
+
"clip_ratio/low_mean": 0.0,
|
| 259 |
+
"clip_ratio/low_min": 0.0,
|
| 260 |
+
"clip_ratio/region_mean": 0.0,
|
| 261 |
+
"completions/clipped_ratio": 0.975,
|
| 262 |
+
"completions/max_length": 256.0,
|
| 263 |
+
"completions/max_terminated_length": 20.4,
|
| 264 |
+
"completions/mean_length": 250.875,
|
| 265 |
+
"completions/mean_terminated_length": 20.4,
|
| 266 |
+
"completions/min_length": 174.0,
|
| 267 |
+
"completions/min_terminated_length": 20.4,
|
| 268 |
+
"entropy": 1.0103468239307403,
|
| 269 |
+
"epoch": 1.0,
|
| 270 |
+
"frac_reward_zero_std": 0.8,
|
| 271 |
+
"grad_norm": 0.0,
|
| 272 |
+
"learning_rate": 4.5100000000000005e-05,
|
| 273 |
+
"loss": -2.2351741790771484e-09,
|
| 274 |
+
"num_tokens": 347364.0,
|
| 275 |
+
"reward": 0.04707500040531158,
|
| 276 |
+
"reward_std": 0.12459058165550232,
|
| 277 |
+
"rewards/dbre_reward/mean": 0.04707500040531158,
|
| 278 |
+
"rewards/dbre_reward/std": 0.12459058165550232,
|
| 279 |
+
"step": 50,
|
| 280 |
+
"step_time": 36.14326306019211
|
| 281 |
+
},
|
| 282 |
+
{
|
| 283 |
+
"clip_ratio/high_max": 0.0,
|
| 284 |
+
"clip_ratio/high_mean": 0.0,
|
| 285 |
+
"clip_ratio/low_mean": 0.0,
|
| 286 |
+
"clip_ratio/low_min": 0.0,
|
| 287 |
+
"clip_ratio/region_mean": 0.0,
|
| 288 |
+
"completions/clipped_ratio": 0.9375,
|
| 289 |
+
"completions/max_length": 256.0,
|
| 290 |
+
"completions/max_terminated_length": 41.8,
|
| 291 |
+
"completions/mean_length": 244.5625,
|
| 292 |
+
"completions/mean_terminated_length": 18.55,
|
| 293 |
+
"completions/min_length": 161.4,
|
| 294 |
+
"completions/min_terminated_length": 7.8,
|
| 295 |
+
"entropy": 0.9930311724543571,
|
| 296 |
+
"epoch": 1.1,
|
| 297 |
+
"frac_reward_zero_std": 0.6,
|
| 298 |
+
"grad_norm": 0.06396484375,
|
| 299 |
+
"learning_rate": 4.46e-05,
|
| 300 |
+
"loss": -0.017940016090869905,
|
| 301 |
+
"num_tokens": 381409.0,
|
| 302 |
+
"reward": 0.06066250056028366,
|
| 303 |
+
"reward_std": 0.2124839812517166,
|
| 304 |
+
"rewards/dbre_reward/mean": 0.06066250056028366,
|
| 305 |
+
"rewards/dbre_reward/std": 0.21248398423194886,
|
| 306 |
+
"step": 55,
|
| 307 |
+
"step_time": 28.94829335878603
|
| 308 |
+
},
|
| 309 |
+
{
|
| 310 |
+
"clip_ratio/high_max": 0.0,
|
| 311 |
+
"clip_ratio/high_mean": 0.0,
|
| 312 |
+
"clip_ratio/low_mean": 0.0,
|
| 313 |
+
"clip_ratio/low_min": 0.0,
|
| 314 |
+
"clip_ratio/region_mean": 0.0,
|
| 315 |
+
"completions/clipped_ratio": 0.9875,
|
| 316 |
+
"completions/max_length": 256.0,
|
| 317 |
+
"completions/max_terminated_length": 9.6,
|
| 318 |
+
"completions/mean_length": 253.4,
|
| 319 |
+
"completions/mean_terminated_length": 9.6,
|
| 320 |
+
"completions/min_length": 214.4,
|
| 321 |
+
"completions/min_terminated_length": 9.6,
|
| 322 |
+
"entropy": 0.8150956228375434,
|
| 323 |
+
"epoch": 1.2,
|
| 324 |
+
"frac_reward_zero_std": 0.5,
|
| 325 |
+
"grad_norm": 0.06103515625,
|
| 326 |
+
"learning_rate": 4.41e-05,
|
| 327 |
+
"loss": -4.470348358154297e-09,
|
| 328 |
+
"num_tokens": 416161.0,
|
| 329 |
+
"reward": 0.11134999990463257,
|
| 330 |
+
"reward_std": 0.2740528523921967,
|
| 331 |
+
"rewards/dbre_reward/mean": 0.11134999990463257,
|
| 332 |
+
"rewards/dbre_reward/std": 0.27405285835266113,
|
| 333 |
+
"step": 60,
|
| 334 |
+
"step_time": 27.705245727201692
|
| 335 |
+
},
|
| 336 |
+
{
|
| 337 |
+
"clip_ratio/high_max": 0.0,
|
| 338 |
+
"clip_ratio/high_mean": 0.0,
|
| 339 |
+
"clip_ratio/low_mean": 0.0,
|
| 340 |
+
"clip_ratio/low_min": 0.0,
|
| 341 |
+
"clip_ratio/region_mean": 0.0,
|
| 342 |
+
"completions/clipped_ratio": 0.975,
|
| 343 |
+
"completions/max_length": 256.0,
|
| 344 |
+
"completions/max_terminated_length": 41.0,
|
| 345 |
+
"completions/mean_length": 252.1625,
|
| 346 |
+
"completions/mean_terminated_length": 41.0,
|
| 347 |
+
"completions/min_length": 194.6,
|
| 348 |
+
"completions/min_terminated_length": 41.0,
|
| 349 |
+
"entropy": 0.9309644259512424,
|
| 350 |
+
"epoch": 1.3,
|
| 351 |
+
"frac_reward_zero_std": 0.7,
|
| 352 |
+
"grad_norm": 0.0,
|
| 353 |
+
"learning_rate": 4.36e-05,
|
| 354 |
+
"loss": -8.940696716308593e-09,
|
| 355 |
+
"num_tokens": 450814.0,
|
| 356 |
+
"reward": 0.0875,
|
| 357 |
+
"reward_std": 0.21124515533447266,
|
| 358 |
+
"rewards/dbre_reward/mean": 0.0875,
|
| 359 |
+
"rewards/dbre_reward/std": 0.21124515533447266,
|
| 360 |
+
"step": 65,
|
| 361 |
+
"step_time": 27.76657635839365
|
| 362 |
+
},
|
| 363 |
+
{
|
| 364 |
+
"clip_ratio/high_max": 0.0,
|
| 365 |
+
"clip_ratio/high_mean": 0.0,
|
| 366 |
+
"clip_ratio/low_mean": 0.0,
|
| 367 |
+
"clip_ratio/low_min": 0.0,
|
| 368 |
+
"clip_ratio/region_mean": 0.0,
|
| 369 |
+
"completions/clipped_ratio": 0.9875,
|
| 370 |
+
"completions/max_length": 256.0,
|
| 371 |
+
"completions/max_terminated_length": 30.4,
|
| 372 |
+
"completions/mean_length": 254.7,
|
| 373 |
+
"completions/mean_terminated_length": 30.4,
|
| 374 |
+
"completions/min_length": 235.2,
|
| 375 |
+
"completions/min_terminated_length": 30.4,
|
| 376 |
+
"entropy": 0.8473479233682155,
|
| 377 |
+
"epoch": 1.4,
|
| 378 |
+
"frac_reward_zero_std": 0.6,
|
| 379 |
+
"grad_norm": 0.05078125,
|
| 380 |
+
"learning_rate": 4.3100000000000004e-05,
|
| 381 |
+
"loss": -2.2351741790771484e-09,
|
| 382 |
+
"num_tokens": 485670.0,
|
| 383 |
+
"reward": 0.07392499968409538,
|
| 384 |
+
"reward_std": 0.1946355789899826,
|
| 385 |
+
"rewards/dbre_reward/mean": 0.07392499968409538,
|
| 386 |
+
"rewards/dbre_reward/std": 0.1946355879306793,
|
| 387 |
+
"step": 70,
|
| 388 |
+
"step_time": 27.701584570185513
|
| 389 |
+
},
|
| 390 |
+
{
|
| 391 |
+
"clip_ratio/high_max": 0.0,
|
| 392 |
+
"clip_ratio/high_mean": 0.0,
|
| 393 |
+
"clip_ratio/low_mean": 0.0,
|
| 394 |
+
"clip_ratio/low_min": 0.0,
|
| 395 |
+
"clip_ratio/region_mean": 0.0,
|
| 396 |
+
"completions/clipped_ratio": 0.975,
|
| 397 |
+
"completions/max_length": 256.0,
|
| 398 |
+
"completions/max_terminated_length": 51.0,
|
| 399 |
+
"completions/mean_length": 252.7875,
|
| 400 |
+
"completions/mean_terminated_length": 51.0,
|
| 401 |
+
"completions/min_length": 204.6,
|
| 402 |
+
"completions/min_terminated_length": 51.0,
|
| 403 |
+
"entropy": 0.8822006396949291,
|
| 404 |
+
"epoch": 1.5,
|
| 405 |
+
"frac_reward_zero_std": 0.4,
|
| 406 |
+
"grad_norm": 0.046875,
|
| 407 |
+
"learning_rate": 4.26e-05,
|
| 408 |
+
"loss": -0.004587128758430481,
|
| 409 |
+
"num_tokens": 520373.0,
|
| 410 |
+
"reward": 0.09866249859333039,
|
| 411 |
+
"reward_std": 0.29479086995124815,
|
| 412 |
+
"rewards/dbre_reward/mean": 0.09866249859333039,
|
| 413 |
+
"rewards/dbre_reward/std": 0.29479087591171266,
|
| 414 |
+
"step": 75,
|
| 415 |
+
"step_time": 27.723117466596886
|
| 416 |
+
},
|
| 417 |
+
{
|
| 418 |
+
"clip_ratio/high_max": 0.0,
|
| 419 |
+
"clip_ratio/high_mean": 0.0,
|
| 420 |
+
"clip_ratio/low_mean": 0.0,
|
| 421 |
+
"clip_ratio/low_min": 0.0,
|
| 422 |
+
"clip_ratio/region_mean": 0.0,
|
| 423 |
+
"completions/clipped_ratio": 1.0,
|
| 424 |
+
"completions/max_length": 256.0,
|
| 425 |
+
"completions/max_terminated_length": 0.0,
|
| 426 |
+
"completions/mean_length": 256.0,
|
| 427 |
+
"completions/mean_terminated_length": 0.0,
|
| 428 |
+
"completions/min_length": 256.0,
|
| 429 |
+
"completions/min_terminated_length": 0.0,
|
| 430 |
+
"entropy": 0.8492010429501533,
|
| 431 |
+
"epoch": 1.6,
|
| 432 |
+
"frac_reward_zero_std": 0.4,
|
| 433 |
+
"grad_norm": 0.052734375,
|
| 434 |
+
"learning_rate": 4.21e-05,
|
| 435 |
+
"loss": -4.470348358154297e-09,
|
| 436 |
+
"num_tokens": 555333.0,
|
| 437 |
+
"reward": 0.13631249964237213,
|
| 438 |
+
"reward_std": 0.3453687012195587,
|
| 439 |
+
"rewards/dbre_reward/mean": 0.13631249964237213,
|
| 440 |
+
"rewards/dbre_reward/std": 0.34536872506141664,
|
| 441 |
+
"step": 80,
|
| 442 |
+
"step_time": 27.74355768300593
|
| 443 |
+
},
|
| 444 |
+
{
|
| 445 |
+
"clip_ratio/high_max": 0.0,
|
| 446 |
+
"clip_ratio/high_mean": 0.0,
|
| 447 |
+
"clip_ratio/low_mean": 0.0,
|
| 448 |
+
"clip_ratio/low_min": 0.0,
|
| 449 |
+
"clip_ratio/region_mean": 0.0,
|
| 450 |
+
"completions/clipped_ratio": 0.9875,
|
| 451 |
+
"completions/max_length": 256.0,
|
| 452 |
+
"completions/max_terminated_length": 23.0,
|
| 453 |
+
"completions/mean_length": 254.2375,
|
| 454 |
+
"completions/mean_terminated_length": 23.0,
|
| 455 |
+
"completions/min_length": 227.8,
|
| 456 |
+
"completions/min_terminated_length": 23.0,
|
| 457 |
+
"entropy": 0.9008926346898078,
|
| 458 |
+
"epoch": 1.7,
|
| 459 |
+
"frac_reward_zero_std": 0.4,
|
| 460 |
+
"grad_norm": 0.053466796875,
|
| 461 |
+
"learning_rate": 4.16e-05,
|
| 462 |
+
"loss": -0.0038484178483486177,
|
| 463 |
+
"num_tokens": 590152.0,
|
| 464 |
+
"reward": 0.13577499985694885,
|
| 465 |
+
"reward_std": 0.3495619535446167,
|
| 466 |
+
"rewards/dbre_reward/mean": 0.13577499985694885,
|
| 467 |
+
"rewards/dbre_reward/std": 0.34956197142601014,
|
| 468 |
+
"step": 85,
|
| 469 |
+
"step_time": 27.809819040997535
|
| 470 |
+
},
|
| 471 |
+
{
|
| 472 |
+
"clip_ratio/high_max": 0.0,
|
| 473 |
+
"clip_ratio/high_mean": 0.0,
|
| 474 |
+
"clip_ratio/low_mean": 0.0,
|
| 475 |
+
"clip_ratio/low_min": 0.0,
|
| 476 |
+
"clip_ratio/region_mean": 0.0,
|
| 477 |
+
"completions/clipped_ratio": 0.975,
|
| 478 |
+
"completions/max_length": 256.0,
|
| 479 |
+
"completions/max_terminated_length": 34.4,
|
| 480 |
+
"completions/mean_length": 253.7375,
|
| 481 |
+
"completions/mean_terminated_length": 33.1,
|
| 482 |
+
"completions/min_length": 236.6,
|
| 483 |
+
"completions/min_terminated_length": 31.8,
|
| 484 |
+
"entropy": 0.975049901008606,
|
| 485 |
+
"epoch": 1.8,
|
| 486 |
+
"frac_reward_zero_std": 0.3,
|
| 487 |
+
"grad_norm": 0.08447265625,
|
| 488 |
+
"learning_rate": 4.11e-05,
|
| 489 |
+
"loss": -0.0023170128464698792,
|
| 490 |
+
"num_tokens": 624931.0,
|
| 491 |
+
"reward": 0.14498749673366546,
|
| 492 |
+
"reward_std": 0.3355918139219284,
|
| 493 |
+
"rewards/dbre_reward/mean": 0.14498749673366546,
|
| 494 |
+
"rewards/dbre_reward/std": 0.3355918198823929,
|
| 495 |
+
"step": 90,
|
| 496 |
+
"step_time": 27.872812610585243
|
| 497 |
+
},
|
| 498 |
+
{
|
| 499 |
+
"clip_ratio/high_max": 0.0,
|
| 500 |
+
"clip_ratio/high_mean": 0.0,
|
| 501 |
+
"clip_ratio/low_mean": 0.0,
|
| 502 |
+
"clip_ratio/low_min": 0.0,
|
| 503 |
+
"clip_ratio/region_mean": 0.0,
|
| 504 |
+
"completions/clipped_ratio": 0.95,
|
| 505 |
+
"completions/max_length": 256.0,
|
| 506 |
+
"completions/max_terminated_length": 75.6,
|
| 507 |
+
"completions/mean_length": 250.775,
|
| 508 |
+
"completions/mean_terminated_length": 60.6,
|
| 509 |
+
"completions/min_length": 199.2,
|
| 510 |
+
"completions/min_terminated_length": 45.6,
|
| 511 |
+
"entropy": 0.9378853186964988,
|
| 512 |
+
"epoch": 1.9,
|
| 513 |
+
"frac_reward_zero_std": 0.4,
|
| 514 |
+
"grad_norm": 0.0576171875,
|
| 515 |
+
"learning_rate": 4.0600000000000004e-05,
|
| 516 |
+
"loss": -0.010608357191085816,
|
| 517 |
+
"num_tokens": 659473.0,
|
| 518 |
+
"reward": 0.13513749986886978,
|
| 519 |
+
"reward_std": 0.3332351267337799,
|
| 520 |
+
"rewards/dbre_reward/mean": 0.13513749986886978,
|
| 521 |
+
"rewards/dbre_reward/std": 0.3332351326942444,
|
| 522 |
+
"step": 95,
|
| 523 |
+
"step_time": 27.71230500699894
|
| 524 |
+
},
|
| 525 |
+
{
|
| 526 |
+
"clip_ratio/high_max": 0.0,
|
| 527 |
+
"clip_ratio/high_mean": 0.0,
|
| 528 |
+
"clip_ratio/low_mean": 0.0,
|
| 529 |
+
"clip_ratio/low_min": 0.0,
|
| 530 |
+
"clip_ratio/region_mean": 0.0,
|
| 531 |
+
"completions/clipped_ratio": 0.975,
|
| 532 |
+
"completions/max_length": 256.0,
|
| 533 |
+
"completions/max_terminated_length": 80.0,
|
| 534 |
+
"completions/mean_length": 254.6,
|
| 535 |
+
"completions/mean_terminated_length": 80.0,
|
| 536 |
+
"completions/min_length": 233.6,
|
| 537 |
+
"completions/min_terminated_length": 80.0,
|
| 538 |
+
"entropy": 0.8716862492263318,
|
| 539 |
+
"epoch": 2.0,
|
| 540 |
+
"frac_reward_zero_std": 0.4,
|
| 541 |
+
"grad_norm": 0.06396484375,
|
| 542 |
+
"learning_rate": 4.0100000000000006e-05,
|
| 543 |
+
"loss": 0.0028070926666259764,
|
| 544 |
+
"num_tokens": 694321.0,
|
| 545 |
+
"reward": 0.13687500059604646,
|
| 546 |
+
"reward_std": 0.33139119744300843,
|
| 547 |
+
"rewards/dbre_reward/mean": 0.13687500059604646,
|
| 548 |
+
"rewards/dbre_reward/std": 0.3313912093639374,
|
| 549 |
+
"step": 100,
|
| 550 |
+
"step_time": 27.905723336405934
|
| 551 |
+
},
|
| 552 |
+
{
|
| 553 |
+
"clip_ratio/high_max": 0.0,
|
| 554 |
+
"clip_ratio/high_mean": 0.0,
|
| 555 |
+
"clip_ratio/low_mean": 0.0,
|
| 556 |
+
"clip_ratio/low_min": 0.0,
|
| 557 |
+
"clip_ratio/region_mean": 0.0,
|
| 558 |
+
"completions/clipped_ratio": 0.95,
|
| 559 |
+
"completions/max_length": 256.0,
|
| 560 |
+
"completions/max_terminated_length": 138.6,
|
| 561 |
+
"completions/mean_length": 251.8625,
|
| 562 |
+
"completions/mean_terminated_length": 138.6,
|
| 563 |
+
"completions/min_length": 189.8,
|
| 564 |
+
"completions/min_terminated_length": 138.6,
|
| 565 |
+
"entropy": 0.9404824480414391,
|
| 566 |
+
"epoch": 2.1,
|
| 567 |
+
"frac_reward_zero_std": 0.5,
|
| 568 |
+
"grad_norm": 0.060791015625,
|
| 569 |
+
"learning_rate": 3.960000000000001e-05,
|
| 570 |
+
"loss": -0.004697377979755402,
|
| 571 |
+
"num_tokens": 728950.0,
|
| 572 |
+
"reward": 0.08519999980926514,
|
| 573 |
+
"reward_std": 0.24869290590286255,
|
| 574 |
+
"rewards/dbre_reward/mean": 0.08519999980926514,
|
| 575 |
+
"rewards/dbre_reward/std": 0.24869290590286255,
|
| 576 |
+
"step": 105,
|
| 577 |
+
"step_time": 27.779890004795742
|
| 578 |
+
},
|
| 579 |
+
{
|
| 580 |
+
"clip_ratio/high_max": 0.0,
|
| 581 |
+
"clip_ratio/high_mean": 0.0,
|
| 582 |
+
"clip_ratio/low_mean": 0.0,
|
| 583 |
+
"clip_ratio/low_min": 0.0,
|
| 584 |
+
"clip_ratio/region_mean": 0.0,
|
| 585 |
+
"completions/clipped_ratio": 0.9625,
|
| 586 |
+
"completions/max_length": 256.0,
|
| 587 |
+
"completions/max_terminated_length": 21.8,
|
| 588 |
+
"completions/mean_length": 248.425,
|
| 589 |
+
"completions/mean_terminated_length": 19.9,
|
| 590 |
+
"completions/min_length": 171.6,
|
| 591 |
+
"completions/min_terminated_length": 18.0,
|
| 592 |
+
"entropy": 1.0435175843536855,
|
| 593 |
+
"epoch": 2.2,
|
| 594 |
+
"frac_reward_zero_std": 0.3,
|
| 595 |
+
"grad_norm": 0.134765625,
|
| 596 |
+
"learning_rate": 3.91e-05,
|
| 597 |
+
"loss": -0.007500007748603821,
|
| 598 |
+
"num_tokens": 763304.0,
|
| 599 |
+
"reward": 0.13616250157356263,
|
| 600 |
+
"reward_std": 0.3243652701377869,
|
| 601 |
+
"rewards/dbre_reward/mean": 0.13616250157356263,
|
| 602 |
+
"rewards/dbre_reward/std": 0.3243652701377869,
|
| 603 |
+
"step": 110,
|
| 604 |
+
"step_time": 27.559503265400416
|
| 605 |
+
},
|
| 606 |
+
{
|
| 607 |
+
"clip_ratio/high_max": 0.0,
|
| 608 |
+
"clip_ratio/high_mean": 0.0,
|
| 609 |
+
"clip_ratio/low_mean": 0.0,
|
| 610 |
+
"clip_ratio/low_min": 0.0,
|
| 611 |
+
"clip_ratio/region_mean": 0.0,
|
| 612 |
+
"completions/clipped_ratio": 0.9875,
|
| 613 |
+
"completions/max_length": 256.0,
|
| 614 |
+
"completions/max_terminated_length": 10.0,
|
| 615 |
+
"completions/mean_length": 253.425,
|
| 616 |
+
"completions/mean_terminated_length": 10.0,
|
| 617 |
+
"completions/min_length": 214.8,
|
| 618 |
+
"completions/min_terminated_length": 10.0,
|
| 619 |
+
"entropy": 0.9914286866784096,
|
| 620 |
+
"epoch": 2.3,
|
| 621 |
+
"frac_reward_zero_std": 0.2,
|
| 622 |
+
"grad_norm": 0.0830078125,
|
| 623 |
+
"learning_rate": 3.86e-05,
|
| 624 |
+
"loss": -0.0037435129284858703,
|
| 625 |
+
"num_tokens": 798058.0,
|
| 626 |
+
"reward": 0.17202500104904175,
|
| 627 |
+
"reward_std": 0.37993268966674804,
|
| 628 |
+
"rewards/dbre_reward/mean": 0.17202500104904175,
|
| 629 |
+
"rewards/dbre_reward/std": 0.37993271350860597,
|
| 630 |
+
"step": 115,
|
| 631 |
+
"step_time": 27.579605548398103
|
| 632 |
+
},
|
| 633 |
+
{
|
| 634 |
+
"clip_ratio/high_max": 0.0,
|
| 635 |
+
"clip_ratio/high_mean": 0.0,
|
| 636 |
+
"clip_ratio/low_mean": 0.0,
|
| 637 |
+
"clip_ratio/low_min": 0.0,
|
| 638 |
+
"clip_ratio/region_mean": 0.0,
|
| 639 |
+
"completions/clipped_ratio": 0.95,
|
| 640 |
+
"completions/max_length": 256.0,
|
| 641 |
+
"completions/max_terminated_length": 128.6,
|
| 642 |
+
"completions/mean_length": 252.5,
|
| 643 |
+
"completions/mean_terminated_length": 117.6,
|
| 644 |
+
"completions/min_length": 209.0,
|
| 645 |
+
"completions/min_terminated_length": 106.6,
|
| 646 |
+
"entropy": 1.0246646717190742,
|
| 647 |
+
"epoch": 2.4,
|
| 648 |
+
"frac_reward_zero_std": 0.2,
|
| 649 |
+
"grad_norm": 0.0927734375,
|
| 650 |
+
"learning_rate": 3.8100000000000005e-05,
|
| 651 |
+
"loss": -0.004933140054345131,
|
| 652 |
+
"num_tokens": 832738.0,
|
| 653 |
+
"reward": 0.18355000019073486,
|
| 654 |
+
"reward_std": 0.36511489748954773,
|
| 655 |
+
"rewards/dbre_reward/mean": 0.18355000019073486,
|
| 656 |
+
"rewards/dbre_reward/std": 0.36511489748954773,
|
| 657 |
+
"step": 120,
|
| 658 |
+
"step_time": 27.58797151680337
|
| 659 |
+
},
|
| 660 |
+
{
|
| 661 |
+
"clip_ratio/high_max": 0.0,
|
| 662 |
+
"clip_ratio/high_mean": 0.0,
|
| 663 |
+
"clip_ratio/low_mean": 0.0,
|
| 664 |
+
"clip_ratio/low_min": 0.0,
|
| 665 |
+
"clip_ratio/region_mean": 0.0,
|
| 666 |
+
"completions/clipped_ratio": 0.9875,
|
| 667 |
+
"completions/max_length": 256.0,
|
| 668 |
+
"completions/max_terminated_length": 5.2,
|
| 669 |
+
"completions/mean_length": 253.125,
|
| 670 |
+
"completions/mean_terminated_length": 5.2,
|
| 671 |
+
"completions/min_length": 210.0,
|
| 672 |
+
"completions/min_terminated_length": 5.2,
|
| 673 |
+
"entropy": 0.9524292007088662,
|
| 674 |
+
"epoch": 2.5,
|
| 675 |
+
"frac_reward_zero_std": 0.3,
|
| 676 |
+
"grad_norm": 0.083984375,
|
| 677 |
+
"learning_rate": 3.76e-05,
|
| 678 |
+
"loss": -0.004205547273159027,
|
| 679 |
+
"num_tokens": 867468.0,
|
| 680 |
+
"reward": 0.22202499732375144,
|
| 681 |
+
"reward_std": 0.3661981552839279,
|
| 682 |
+
"rewards/dbre_reward/mean": 0.22202499732375144,
|
| 683 |
+
"rewards/dbre_reward/std": 0.3661981612443924,
|
| 684 |
+
"step": 125,
|
| 685 |
+
"step_time": 27.638399406394456
|
| 686 |
+
},
|
| 687 |
+
{
|
| 688 |
+
"clip_ratio/high_max": 0.0,
|
| 689 |
+
"clip_ratio/high_mean": 0.0,
|
| 690 |
+
"clip_ratio/low_mean": 0.0,
|
| 691 |
+
"clip_ratio/low_min": 0.0,
|
| 692 |
+
"clip_ratio/region_mean": 0.0,
|
| 693 |
+
"completions/clipped_ratio": 0.9875,
|
| 694 |
+
"completions/max_length": 256.0,
|
| 695 |
+
"completions/max_terminated_length": 34.4,
|
| 696 |
+
"completions/mean_length": 254.95,
|
| 697 |
+
"completions/mean_terminated_length": 34.4,
|
| 698 |
+
"completions/min_length": 239.2,
|
| 699 |
+
"completions/min_terminated_length": 34.4,
|
| 700 |
+
"entropy": 0.9268762037158013,
|
| 701 |
+
"epoch": 2.6,
|
| 702 |
+
"frac_reward_zero_std": 0.4,
|
| 703 |
+
"grad_norm": 0.060546875,
|
| 704 |
+
"learning_rate": 3.71e-05,
|
| 705 |
+
"loss": -0.002260996401309967,
|
| 706 |
+
"num_tokens": 902344.0,
|
| 707 |
+
"reward": 0.19577499628067016,
|
| 708 |
+
"reward_std": 0.3975414574146271,
|
| 709 |
+
"rewards/dbre_reward/mean": 0.19577499628067016,
|
| 710 |
+
"rewards/dbre_reward/std": 0.397541481256485,
|
| 711 |
+
"step": 130,
|
| 712 |
+
"step_time": 27.68072868139425
|
| 713 |
+
},
|
| 714 |
+
{
|
| 715 |
+
"clip_ratio/high_max": 0.0,
|
| 716 |
+
"clip_ratio/high_mean": 0.0,
|
| 717 |
+
"clip_ratio/low_mean": 0.0,
|
| 718 |
+
"clip_ratio/low_min": 0.0,
|
| 719 |
+
"clip_ratio/region_mean": 0.0,
|
| 720 |
+
"completions/clipped_ratio": 0.9875,
|
| 721 |
+
"completions/max_length": 256.0,
|
| 722 |
+
"completions/max_terminated_length": 26.6,
|
| 723 |
+
"completions/mean_length": 254.4625,
|
| 724 |
+
"completions/mean_terminated_length": 26.6,
|
| 725 |
+
"completions/min_length": 231.4,
|
| 726 |
+
"completions/min_terminated_length": 26.6,
|
| 727 |
+
"entropy": 0.9508431695401669,
|
| 728 |
+
"epoch": 2.7,
|
| 729 |
+
"frac_reward_zero_std": 0.1,
|
| 730 |
+
"grad_norm": 0.1025390625,
|
| 731 |
+
"learning_rate": 3.66e-05,
|
| 732 |
+
"loss": -0.004485464096069336,
|
| 733 |
+
"num_tokens": 937181.0,
|
| 734 |
+
"reward": 0.1941875010728836,
|
| 735 |
+
"reward_std": 0.39680722951889036,
|
| 736 |
+
"rewards/dbre_reward/mean": 0.1941875010728836,
|
| 737 |
+
"rewards/dbre_reward/std": 0.39680724740028384,
|
| 738 |
+
"step": 135,
|
| 739 |
+
"step_time": 27.79836461079831
|
| 740 |
+
},
|
| 741 |
+
{
|
| 742 |
+
"clip_ratio/high_max": 0.0,
|
| 743 |
+
"clip_ratio/high_mean": 0.0,
|
| 744 |
+
"clip_ratio/low_mean": 0.0,
|
| 745 |
+
"clip_ratio/low_min": 0.0,
|
| 746 |
+
"clip_ratio/region_mean": 0.0,
|
| 747 |
+
"completions/clipped_ratio": 1.0,
|
| 748 |
+
"completions/max_length": 256.0,
|
| 749 |
+
"completions/max_terminated_length": 0.0,
|
| 750 |
+
"completions/mean_length": 256.0,
|
| 751 |
+
"completions/mean_terminated_length": 0.0,
|
| 752 |
+
"completions/min_length": 256.0,
|
| 753 |
+
"completions/min_terminated_length": 0.0,
|
| 754 |
+
"entropy": 0.9166437476873398,
|
| 755 |
+
"epoch": 2.8,
|
| 756 |
+
"frac_reward_zero_std": 0.1,
|
| 757 |
+
"grad_norm": 0.078125,
|
| 758 |
+
"learning_rate": 3.61e-05,
|
| 759 |
+
"loss": 2.9802322387695314e-09,
|
| 760 |
+
"num_tokens": 972141.0,
|
| 761 |
+
"reward": 0.1720750018954277,
|
| 762 |
+
"reward_std": 0.3809880971908569,
|
| 763 |
+
"rewards/dbre_reward/mean": 0.1720750018954277,
|
| 764 |
+
"rewards/dbre_reward/std": 0.3809881091117859,
|
| 765 |
+
"step": 140,
|
| 766 |
+
"step_time": 27.89382164280105
|
| 767 |
+
},
|
| 768 |
+
{
|
| 769 |
+
"clip_ratio/high_max": 0.0,
|
| 770 |
+
"clip_ratio/high_mean": 0.0,
|
| 771 |
+
"clip_ratio/low_mean": 0.0,
|
| 772 |
+
"clip_ratio/low_min": 0.0,
|
| 773 |
+
"clip_ratio/region_mean": 0.0,
|
| 774 |
+
"completions/clipped_ratio": 0.95,
|
| 775 |
+
"completions/max_length": 256.0,
|
| 776 |
+
"completions/max_terminated_length": 117.2,
|
| 777 |
+
"completions/mean_length": 252.3125,
|
| 778 |
+
"completions/mean_terminated_length": 108.2,
|
| 779 |
+
"completions/min_length": 201.6,
|
| 780 |
+
"completions/min_terminated_length": 99.2,
|
| 781 |
+
"entropy": 1.039070624113083,
|
| 782 |
+
"epoch": 2.9,
|
| 783 |
+
"frac_reward_zero_std": 0.2,
|
| 784 |
+
"grad_norm": 0.08740234375,
|
| 785 |
+
"learning_rate": 3.56e-05,
|
| 786 |
+
"loss": 0.003438304364681244,
|
| 787 |
+
"num_tokens": 1006806.0,
|
| 788 |
+
"reward": 0.2208999961614609,
|
| 789 |
+
"reward_std": 0.401767635345459,
|
| 790 |
+
"rewards/dbre_reward/mean": 0.2208999961614609,
|
| 791 |
+
"rewards/dbre_reward/std": 0.4017676472663879,
|
| 792 |
+
"step": 145,
|
| 793 |
+
"step_time": 27.960635445202932
|
| 794 |
+
},
|
| 795 |
+
{
|
| 796 |
+
"clip_ratio/high_max": 0.0,
|
| 797 |
+
"clip_ratio/high_mean": 0.0,
|
| 798 |
+
"clip_ratio/low_mean": 0.0,
|
| 799 |
+
"clip_ratio/low_min": 0.0,
|
| 800 |
+
"clip_ratio/region_mean": 0.0,
|
| 801 |
+
"completions/clipped_ratio": 0.975,
|
| 802 |
+
"completions/max_length": 256.0,
|
| 803 |
+
"completions/max_terminated_length": 40.2,
|
| 804 |
+
"completions/mean_length": 253.0625,
|
| 805 |
+
"completions/mean_terminated_length": 27.7,
|
| 806 |
+
"completions/min_length": 220.0,
|
| 807 |
+
"completions/min_terminated_length": 15.2,
|
| 808 |
+
"entropy": 0.9523506201803684,
|
| 809 |
+
"epoch": 3.0,
|
| 810 |
+
"frac_reward_zero_std": 0.1,
|
| 811 |
+
"grad_norm": 0.0966796875,
|
| 812 |
+
"learning_rate": 3.51e-05,
|
| 813 |
+
"loss": -0.006567706167697906,
|
| 814 |
+
"num_tokens": 1041531.0,
|
| 815 |
+
"reward": 0.254237499833107,
|
| 816 |
+
"reward_std": 0.4319828271865845,
|
| 817 |
+
"rewards/dbre_reward/mean": 0.254237499833107,
|
| 818 |
+
"rewards/dbre_reward/std": 0.4319828271865845,
|
| 819 |
+
"step": 150,
|
| 820 |
+
"step_time": 33.83420682080032
|
| 821 |
+
}
|
| 822 |
+
],
|
| 823 |
+
"logging_steps": 5,
|
| 824 |
+
"max_steps": 500,
|
| 825 |
+
"num_input_tokens_seen": 1041531,
|
| 826 |
+
"num_train_epochs": 10,
|
| 827 |
+
"save_steps": 50,
|
| 828 |
+
"stateful_callbacks": {
|
| 829 |
+
"TrainerControl": {
|
| 830 |
+
"args": {
|
| 831 |
+
"should_epoch_stop": false,
|
| 832 |
+
"should_evaluate": false,
|
| 833 |
+
"should_log": false,
|
| 834 |
+
"should_save": true,
|
| 835 |
+
"should_training_stop": false
|
| 836 |
+
},
|
| 837 |
+
"attributes": {}
|
| 838 |
+
}
|
| 839 |
+
},
|
| 840 |
+
"total_flos": 0.0,
|
| 841 |
+
"train_batch_size": 2,
|
| 842 |
+
"trial_name": null,
|
| 843 |
+
"trial_params": null
|
| 844 |
+
}
|
grpo_dbre/checkpoint-150/training_args.bin
ADDED
|
Binary file (7.12 kB). View file
|
|
|
grpo_dbre/checkpoint-200/README.md
ADDED
|
@@ -0,0 +1,209 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
base_model: Qwen/Qwen2.5-Coder-1.5B-Instruct
|
| 3 |
+
library_name: peft
|
| 4 |
+
pipeline_tag: text-generation
|
| 5 |
+
tags:
|
| 6 |
+
- base_model:adapter:Qwen/Qwen2.5-Coder-1.5B-Instruct
|
| 7 |
+
- grpo
|
| 8 |
+
- lora
|
| 9 |
+
- transformers
|
| 10 |
+
- trl
|
| 11 |
+
---
|
| 12 |
+
|
| 13 |
+
# Model Card for Model ID
|
| 14 |
+
|
| 15 |
+
<!-- Provide a quick summary of what the model is/does. -->
|
| 16 |
+
|
| 17 |
+
|
| 18 |
+
|
| 19 |
+
## Model Details
|
| 20 |
+
|
| 21 |
+
### Model Description
|
| 22 |
+
|
| 23 |
+
<!-- Provide a longer summary of what this model is. -->
|
| 24 |
+
|
| 25 |
+
|
| 26 |
+
|
| 27 |
+
- **Developed by:** [More Information Needed]
|
| 28 |
+
- **Funded by [optional]:** [More Information Needed]
|
| 29 |
+
- **Shared by [optional]:** [More Information Needed]
|
| 30 |
+
- **Model type:** [More Information Needed]
|
| 31 |
+
- **Language(s) (NLP):** [More Information Needed]
|
| 32 |
+
- **License:** [More Information Needed]
|
| 33 |
+
- **Finetuned from model [optional]:** [More Information Needed]
|
| 34 |
+
|
| 35 |
+
### Model Sources [optional]
|
| 36 |
+
|
| 37 |
+
<!-- Provide the basic links for the model. -->
|
| 38 |
+
|
| 39 |
+
- **Repository:** [More Information Needed]
|
| 40 |
+
- **Paper [optional]:** [More Information Needed]
|
| 41 |
+
- **Demo [optional]:** [More Information Needed]
|
| 42 |
+
|
| 43 |
+
## Uses
|
| 44 |
+
|
| 45 |
+
<!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
|
| 46 |
+
|
| 47 |
+
### Direct Use
|
| 48 |
+
|
| 49 |
+
<!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
|
| 50 |
+
|
| 51 |
+
[More Information Needed]
|
| 52 |
+
|
| 53 |
+
### Downstream Use [optional]
|
| 54 |
+
|
| 55 |
+
<!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
|
| 56 |
+
|
| 57 |
+
[More Information Needed]
|
| 58 |
+
|
| 59 |
+
### Out-of-Scope Use
|
| 60 |
+
|
| 61 |
+
<!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
|
| 62 |
+
|
| 63 |
+
[More Information Needed]
|
| 64 |
+
|
| 65 |
+
## Bias, Risks, and Limitations
|
| 66 |
+
|
| 67 |
+
<!-- This section is meant to convey both technical and sociotechnical limitations. -->
|
| 68 |
+
|
| 69 |
+
[More Information Needed]
|
| 70 |
+
|
| 71 |
+
### Recommendations
|
| 72 |
+
|
| 73 |
+
<!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
|
| 74 |
+
|
| 75 |
+
Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
|
| 76 |
+
|
| 77 |
+
## How to Get Started with the Model
|
| 78 |
+
|
| 79 |
+
Use the code below to get started with the model.
|
| 80 |
+
|
| 81 |
+
[More Information Needed]
|
| 82 |
+
|
| 83 |
+
## Training Details
|
| 84 |
+
|
| 85 |
+
### Training Data
|
| 86 |
+
|
| 87 |
+
<!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
|
| 88 |
+
|
| 89 |
+
[More Information Needed]
|
| 90 |
+
|
| 91 |
+
### Training Procedure
|
| 92 |
+
|
| 93 |
+
<!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
|
| 94 |
+
|
| 95 |
+
#### Preprocessing [optional]
|
| 96 |
+
|
| 97 |
+
[More Information Needed]
|
| 98 |
+
|
| 99 |
+
|
| 100 |
+
#### Training Hyperparameters
|
| 101 |
+
|
| 102 |
+
- **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
|
| 103 |
+
|
| 104 |
+
#### Speeds, Sizes, Times [optional]
|
| 105 |
+
|
| 106 |
+
<!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
|
| 107 |
+
|
| 108 |
+
[More Information Needed]
|
| 109 |
+
|
| 110 |
+
## Evaluation
|
| 111 |
+
|
| 112 |
+
<!-- This section describes the evaluation protocols and provides the results. -->
|
| 113 |
+
|
| 114 |
+
### Testing Data, Factors & Metrics
|
| 115 |
+
|
| 116 |
+
#### Testing Data
|
| 117 |
+
|
| 118 |
+
<!-- This should link to a Dataset Card if possible. -->
|
| 119 |
+
|
| 120 |
+
[More Information Needed]
|
| 121 |
+
|
| 122 |
+
#### Factors
|
| 123 |
+
|
| 124 |
+
<!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
|
| 125 |
+
|
| 126 |
+
[More Information Needed]
|
| 127 |
+
|
| 128 |
+
#### Metrics
|
| 129 |
+
|
| 130 |
+
<!-- These are the evaluation metrics being used, ideally with a description of why. -->
|
| 131 |
+
|
| 132 |
+
[More Information Needed]
|
| 133 |
+
|
| 134 |
+
### Results
|
| 135 |
+
|
| 136 |
+
[More Information Needed]
|
| 137 |
+
|
| 138 |
+
#### Summary
|
| 139 |
+
|
| 140 |
+
|
| 141 |
+
|
| 142 |
+
## Model Examination [optional]
|
| 143 |
+
|
| 144 |
+
<!-- Relevant interpretability work for the model goes here -->
|
| 145 |
+
|
| 146 |
+
[More Information Needed]
|
| 147 |
+
|
| 148 |
+
## Environmental Impact
|
| 149 |
+
|
| 150 |
+
<!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
|
| 151 |
+
|
| 152 |
+
Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
|
| 153 |
+
|
| 154 |
+
- **Hardware Type:** [More Information Needed]
|
| 155 |
+
- **Hours used:** [More Information Needed]
|
| 156 |
+
- **Cloud Provider:** [More Information Needed]
|
| 157 |
+
- **Compute Region:** [More Information Needed]
|
| 158 |
+
- **Carbon Emitted:** [More Information Needed]
|
| 159 |
+
|
| 160 |
+
## Technical Specifications [optional]
|
| 161 |
+
|
| 162 |
+
### Model Architecture and Objective
|
| 163 |
+
|
| 164 |
+
[More Information Needed]
|
| 165 |
+
|
| 166 |
+
### Compute Infrastructure
|
| 167 |
+
|
| 168 |
+
[More Information Needed]
|
| 169 |
+
|
| 170 |
+
#### Hardware
|
| 171 |
+
|
| 172 |
+
[More Information Needed]
|
| 173 |
+
|
| 174 |
+
#### Software
|
| 175 |
+
|
| 176 |
+
[More Information Needed]
|
| 177 |
+
|
| 178 |
+
## Citation [optional]
|
| 179 |
+
|
| 180 |
+
<!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
|
| 181 |
+
|
| 182 |
+
**BibTeX:**
|
| 183 |
+
|
| 184 |
+
[More Information Needed]
|
| 185 |
+
|
| 186 |
+
**APA:**
|
| 187 |
+
|
| 188 |
+
[More Information Needed]
|
| 189 |
+
|
| 190 |
+
## Glossary [optional]
|
| 191 |
+
|
| 192 |
+
<!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
|
| 193 |
+
|
| 194 |
+
[More Information Needed]
|
| 195 |
+
|
| 196 |
+
## More Information [optional]
|
| 197 |
+
|
| 198 |
+
[More Information Needed]
|
| 199 |
+
|
| 200 |
+
## Model Card Authors [optional]
|
| 201 |
+
|
| 202 |
+
[More Information Needed]
|
| 203 |
+
|
| 204 |
+
## Model Card Contact
|
| 205 |
+
|
| 206 |
+
[More Information Needed]
|
| 207 |
+
### Framework versions
|
| 208 |
+
|
| 209 |
+
- PEFT 0.19.1
|
grpo_dbre/checkpoint-200/adapter_config.json
ADDED
|
@@ -0,0 +1,45 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"alora_invocation_tokens": null,
|
| 3 |
+
"alpha_pattern": {},
|
| 4 |
+
"arrow_config": null,
|
| 5 |
+
"auto_mapping": null,
|
| 6 |
+
"base_model_name_or_path": "Qwen/Qwen2.5-Coder-1.5B-Instruct",
|
| 7 |
+
"bias": "none",
|
| 8 |
+
"corda_config": null,
|
| 9 |
+
"ensure_weight_tying": false,
|
| 10 |
+
"eva_config": null,
|
| 11 |
+
"exclude_modules": null,
|
| 12 |
+
"fan_in_fan_out": false,
|
| 13 |
+
"inference_mode": true,
|
| 14 |
+
"init_lora_weights": true,
|
| 15 |
+
"layer_replication": null,
|
| 16 |
+
"layers_pattern": null,
|
| 17 |
+
"layers_to_transform": null,
|
| 18 |
+
"loftq_config": {},
|
| 19 |
+
"lora_alpha": 16,
|
| 20 |
+
"lora_bias": false,
|
| 21 |
+
"lora_dropout": 0,
|
| 22 |
+
"lora_ga_config": null,
|
| 23 |
+
"megatron_config": null,
|
| 24 |
+
"megatron_core": "megatron.core",
|
| 25 |
+
"modules_to_save": null,
|
| 26 |
+
"peft_type": "LORA",
|
| 27 |
+
"peft_version": "0.19.1",
|
| 28 |
+
"qalora_group_size": 16,
|
| 29 |
+
"r": 16,
|
| 30 |
+
"rank_pattern": {},
|
| 31 |
+
"revision": null,
|
| 32 |
+
"target_modules": [
|
| 33 |
+
"q_proj",
|
| 34 |
+
"k_proj",
|
| 35 |
+
"o_proj",
|
| 36 |
+
"v_proj"
|
| 37 |
+
],
|
| 38 |
+
"target_parameters": null,
|
| 39 |
+
"task_type": "CAUSAL_LM",
|
| 40 |
+
"trainable_token_indices": null,
|
| 41 |
+
"use_bdlora": null,
|
| 42 |
+
"use_dora": false,
|
| 43 |
+
"use_qalora": false,
|
| 44 |
+
"use_rslora": false
|
| 45 |
+
}
|
grpo_dbre/checkpoint-200/chat_template.jinja
ADDED
|
@@ -0,0 +1,54 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{%- if tools %}
|
| 2 |
+
{{- '<|im_start|>system\n' }}
|
| 3 |
+
{%- if messages[0]['role'] == 'system' %}
|
| 4 |
+
{{- messages[0]['content'] }}
|
| 5 |
+
{%- else %}
|
| 6 |
+
{{- 'You are Qwen, created by Alibaba Cloud. You are a helpful assistant.' }}
|
| 7 |
+
{%- endif %}
|
| 8 |
+
{{- "\n\n# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>" }}
|
| 9 |
+
{%- for tool in tools %}
|
| 10 |
+
{{- "\n" }}
|
| 11 |
+
{{- tool | tojson }}
|
| 12 |
+
{%- endfor %}
|
| 13 |
+
{{- "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call><|im_end|>\n" }}
|
| 14 |
+
{%- else %}
|
| 15 |
+
{%- if messages[0]['role'] == 'system' %}
|
| 16 |
+
{{- '<|im_start|>system\n' + messages[0]['content'] + '<|im_end|>\n' }}
|
| 17 |
+
{%- else %}
|
| 18 |
+
{{- '<|im_start|>system\nYou are Qwen, created by Alibaba Cloud. You are a helpful assistant.<|im_end|>\n' }}
|
| 19 |
+
{%- endif %}
|
| 20 |
+
{%- endif %}
|
| 21 |
+
{%- for message in messages %}
|
| 22 |
+
{%- if (message.role == "user") or (message.role == "system" and not loop.first) or (message.role == "assistant" and not message.tool_calls) %}
|
| 23 |
+
{{- '<|im_start|>' + message.role + '\n' + message.content + '<|im_end|>' + '\n' }}
|
| 24 |
+
{%- elif message.role == "assistant" %}
|
| 25 |
+
{{- '<|im_start|>' + message.role }}
|
| 26 |
+
{%- if message.content %}
|
| 27 |
+
{{- '\n' + message.content }}
|
| 28 |
+
{%- endif %}
|
| 29 |
+
{%- for tool_call in message.tool_calls %}
|
| 30 |
+
{%- if tool_call.function is defined %}
|
| 31 |
+
{%- set tool_call = tool_call.function %}
|
| 32 |
+
{%- endif %}
|
| 33 |
+
{{- '\n<tool_call>\n{"name": "' }}
|
| 34 |
+
{{- tool_call.name }}
|
| 35 |
+
{{- '", "arguments": ' }}
|
| 36 |
+
{{- tool_call.arguments | tojson }}
|
| 37 |
+
{{- '}\n</tool_call>' }}
|
| 38 |
+
{%- endfor %}
|
| 39 |
+
{{- '<|im_end|>\n' }}
|
| 40 |
+
{%- elif message.role == "tool" %}
|
| 41 |
+
{%- if (loop.index0 == 0) or (messages[loop.index0 - 1].role != "tool") %}
|
| 42 |
+
{{- '<|im_start|>user' }}
|
| 43 |
+
{%- endif %}
|
| 44 |
+
{{- '\n<tool_response>\n' }}
|
| 45 |
+
{{- message.content }}
|
| 46 |
+
{{- '\n</tool_response>' }}
|
| 47 |
+
{%- if loop.last or (messages[loop.index0 + 1].role != "tool") %}
|
| 48 |
+
{{- '<|im_end|>\n' }}
|
| 49 |
+
{%- endif %}
|
| 50 |
+
{%- endif %}
|
| 51 |
+
{%- endfor %}
|
| 52 |
+
{%- if add_generation_prompt %}
|
| 53 |
+
{{- '<|im_start|>assistant\n' }}
|
| 54 |
+
{%- endif %}
|
grpo_dbre/checkpoint-200/rng_state.pth
ADDED
|
Binary file (14.6 kB). View file
|
|
|
grpo_dbre/checkpoint-200/scheduler.pt
ADDED
|
Binary file (1.47 kB). View file
|
|
|
grpo_dbre/checkpoint-200/tokenizer_config.json
ADDED
|
@@ -0,0 +1,30 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"add_prefix_space": false,
|
| 3 |
+
"backend": "tokenizers",
|
| 4 |
+
"bos_token": null,
|
| 5 |
+
"clean_up_tokenization_spaces": false,
|
| 6 |
+
"eos_token": "<|im_end|>",
|
| 7 |
+
"errors": "replace",
|
| 8 |
+
"extra_special_tokens": [
|
| 9 |
+
"<|im_start|>",
|
| 10 |
+
"<|im_end|>",
|
| 11 |
+
"<|object_ref_start|>",
|
| 12 |
+
"<|object_ref_end|>",
|
| 13 |
+
"<|box_start|>",
|
| 14 |
+
"<|box_end|>",
|
| 15 |
+
"<|quad_start|>",
|
| 16 |
+
"<|quad_end|>",
|
| 17 |
+
"<|vision_start|>",
|
| 18 |
+
"<|vision_end|>",
|
| 19 |
+
"<|vision_pad|>",
|
| 20 |
+
"<|image_pad|>",
|
| 21 |
+
"<|video_pad|>"
|
| 22 |
+
],
|
| 23 |
+
"is_local": false,
|
| 24 |
+
"local_files_only": false,
|
| 25 |
+
"model_max_length": 32768,
|
| 26 |
+
"pad_token": "<|im_end|>",
|
| 27 |
+
"split_special_tokens": false,
|
| 28 |
+
"tokenizer_class": "Qwen2Tokenizer",
|
| 29 |
+
"unk_token": null
|
| 30 |
+
}
|
grpo_dbre/checkpoint-200/trainer_state.json
ADDED
|
@@ -0,0 +1,1114 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"best_global_step": null,
|
| 3 |
+
"best_metric": null,
|
| 4 |
+
"best_model_checkpoint": null,
|
| 5 |
+
"epoch": 4.0,
|
| 6 |
+
"eval_steps": 500,
|
| 7 |
+
"global_step": 200,
|
| 8 |
+
"is_hyper_param_search": false,
|
| 9 |
+
"is_local_process_zero": true,
|
| 10 |
+
"is_world_process_zero": true,
|
| 11 |
+
"log_history": [
|
| 12 |
+
{
|
| 13 |
+
"clip_ratio/high_max": 0.0,
|
| 14 |
+
"clip_ratio/high_mean": 0.0,
|
| 15 |
+
"clip_ratio/low_mean": 0.0,
|
| 16 |
+
"clip_ratio/low_min": 0.0,
|
| 17 |
+
"clip_ratio/region_mean": 0.0,
|
| 18 |
+
"completions/clipped_ratio": 0.9875,
|
| 19 |
+
"completions/max_length": 256.0,
|
| 20 |
+
"completions/max_terminated_length": 19.6,
|
| 21 |
+
"completions/mean_length": 254.025,
|
| 22 |
+
"completions/mean_terminated_length": 19.6,
|
| 23 |
+
"completions/min_length": 224.4,
|
| 24 |
+
"completions/min_terminated_length": 19.6,
|
| 25 |
+
"entropy": 0.903753462433815,
|
| 26 |
+
"epoch": 0.1,
|
| 27 |
+
"frac_reward_zero_std": 0.8,
|
| 28 |
+
"grad_norm": 0.046875,
|
| 29 |
+
"learning_rate": 4.96e-05,
|
| 30 |
+
"loss": -2.2351741790771484e-09,
|
| 31 |
+
"num_tokens": 34802.0,
|
| 32 |
+
"reward": 0.025,
|
| 33 |
+
"reward_std": 0.1,
|
| 34 |
+
"rewards/dbre_reward/mean": 0.025,
|
| 35 |
+
"rewards/dbre_reward/std": 0.1,
|
| 36 |
+
"step": 5,
|
| 37 |
+
"step_time": 27.396174477002933
|
| 38 |
+
},
|
| 39 |
+
{
|
| 40 |
+
"clip_ratio/high_max": 0.0,
|
| 41 |
+
"clip_ratio/high_mean": 0.0,
|
| 42 |
+
"clip_ratio/low_mean": 0.0,
|
| 43 |
+
"clip_ratio/low_min": 0.0,
|
| 44 |
+
"clip_ratio/region_mean": 0.0,
|
| 45 |
+
"completions/clipped_ratio": 0.975,
|
| 46 |
+
"completions/max_length": 256.0,
|
| 47 |
+
"completions/max_terminated_length": 48.4,
|
| 48 |
+
"completions/mean_length": 252.625,
|
| 49 |
+
"completions/mean_terminated_length": 48.4,
|
| 50 |
+
"completions/min_length": 202.0,
|
| 51 |
+
"completions/min_terminated_length": 48.4,
|
| 52 |
+
"entropy": 0.9422640666365624,
|
| 53 |
+
"epoch": 0.2,
|
| 54 |
+
"frac_reward_zero_std": 0.5,
|
| 55 |
+
"grad_norm": 0.0673828125,
|
| 56 |
+
"learning_rate": 4.91e-05,
|
| 57 |
+
"loss": -5.960464477539063e-09,
|
| 58 |
+
"num_tokens": 69492.0,
|
| 59 |
+
"reward": 0.09868749976158142,
|
| 60 |
+
"reward_std": 0.25848535895347596,
|
| 61 |
+
"rewards/dbre_reward/mean": 0.09868749976158142,
|
| 62 |
+
"rewards/dbre_reward/std": 0.2584853649139404,
|
| 63 |
+
"step": 10,
|
| 64 |
+
"step_time": 27.675077842207976
|
| 65 |
+
},
|
| 66 |
+
{
|
| 67 |
+
"clip_ratio/high_max": 0.0,
|
| 68 |
+
"clip_ratio/high_mean": 0.0,
|
| 69 |
+
"clip_ratio/low_mean": 0.0,
|
| 70 |
+
"clip_ratio/low_min": 0.0,
|
| 71 |
+
"clip_ratio/region_mean": 0.0,
|
| 72 |
+
"completions/clipped_ratio": 0.975,
|
| 73 |
+
"completions/max_length": 256.0,
|
| 74 |
+
"completions/max_terminated_length": 79.0,
|
| 75 |
+
"completions/mean_length": 254.5375,
|
| 76 |
+
"completions/mean_terminated_length": 79.0,
|
| 77 |
+
"completions/min_length": 232.6,
|
| 78 |
+
"completions/min_terminated_length": 79.0,
|
| 79 |
+
"entropy": 0.9289400212466716,
|
| 80 |
+
"epoch": 0.3,
|
| 81 |
+
"frac_reward_zero_std": 0.5,
|
| 82 |
+
"grad_norm": 0.059326171875,
|
| 83 |
+
"learning_rate": 4.86e-05,
|
| 84 |
+
"loss": -0.0020709306001663206,
|
| 85 |
+
"num_tokens": 104335.0,
|
| 86 |
+
"reward": 0.11038749814033508,
|
| 87 |
+
"reward_std": 0.27505697011947633,
|
| 88 |
+
"rewards/dbre_reward/mean": 0.11038749814033508,
|
| 89 |
+
"rewards/dbre_reward/std": 0.2750569820404053,
|
| 90 |
+
"step": 15,
|
| 91 |
+
"step_time": 27.678717108402633
|
| 92 |
+
},
|
| 93 |
+
{
|
| 94 |
+
"clip_ratio/high_max": 0.0,
|
| 95 |
+
"clip_ratio/high_mean": 0.0,
|
| 96 |
+
"clip_ratio/low_mean": 0.0,
|
| 97 |
+
"clip_ratio/low_min": 0.0,
|
| 98 |
+
"clip_ratio/region_mean": 0.0,
|
| 99 |
+
"completions/clipped_ratio": 0.9625,
|
| 100 |
+
"completions/max_length": 256.0,
|
| 101 |
+
"completions/max_terminated_length": 61.4,
|
| 102 |
+
"completions/mean_length": 250.2375,
|
| 103 |
+
"completions/mean_terminated_length": 61.4,
|
| 104 |
+
"completions/min_length": 163.8,
|
| 105 |
+
"completions/min_terminated_length": 61.4,
|
| 106 |
+
"entropy": 0.8890757068991662,
|
| 107 |
+
"epoch": 0.4,
|
| 108 |
+
"frac_reward_zero_std": 0.3,
|
| 109 |
+
"grad_norm": 0.052001953125,
|
| 110 |
+
"learning_rate": 4.8100000000000004e-05,
|
| 111 |
+
"loss": -0.00669153705239296,
|
| 112 |
+
"num_tokens": 138834.0,
|
| 113 |
+
"reward": 0.11190000027418137,
|
| 114 |
+
"reward_std": 0.3216355323791504,
|
| 115 |
+
"rewards/dbre_reward/mean": 0.11190000027418137,
|
| 116 |
+
"rewards/dbre_reward/std": 0.3216355502605438,
|
| 117 |
+
"step": 20,
|
| 118 |
+
"step_time": 27.725935825207852
|
| 119 |
+
},
|
| 120 |
+
{
|
| 121 |
+
"clip_ratio/high_max": 0.0,
|
| 122 |
+
"clip_ratio/high_mean": 0.0,
|
| 123 |
+
"clip_ratio/low_mean": 0.0,
|
| 124 |
+
"clip_ratio/low_min": 0.0,
|
| 125 |
+
"clip_ratio/region_mean": 0.0,
|
| 126 |
+
"completions/clipped_ratio": 0.975,
|
| 127 |
+
"completions/max_length": 256.0,
|
| 128 |
+
"completions/max_terminated_length": 78.2,
|
| 129 |
+
"completions/mean_length": 254.4875,
|
| 130 |
+
"completions/mean_terminated_length": 78.2,
|
| 131 |
+
"completions/min_length": 231.8,
|
| 132 |
+
"completions/min_terminated_length": 78.2,
|
| 133 |
+
"entropy": 0.906505486369133,
|
| 134 |
+
"epoch": 0.5,
|
| 135 |
+
"frac_reward_zero_std": 0.7,
|
| 136 |
+
"grad_norm": 0.0,
|
| 137 |
+
"learning_rate": 4.76e-05,
|
| 138 |
+
"loss": 0.00735630989074707,
|
| 139 |
+
"num_tokens": 173673.0,
|
| 140 |
+
"reward": 0.0375,
|
| 141 |
+
"reward_std": 0.15,
|
| 142 |
+
"rewards/dbre_reward/mean": 0.0375,
|
| 143 |
+
"rewards/dbre_reward/std": 0.15,
|
| 144 |
+
"step": 25,
|
| 145 |
+
"step_time": 27.84102148480888
|
| 146 |
+
},
|
| 147 |
+
{
|
| 148 |
+
"clip_ratio/high_max": 0.0,
|
| 149 |
+
"clip_ratio/high_mean": 0.0,
|
| 150 |
+
"clip_ratio/low_mean": 0.0,
|
| 151 |
+
"clip_ratio/low_min": 0.0,
|
| 152 |
+
"clip_ratio/region_mean": 0.0,
|
| 153 |
+
"completions/clipped_ratio": 0.9625,
|
| 154 |
+
"completions/max_length": 256.0,
|
| 155 |
+
"completions/max_terminated_length": 69.0,
|
| 156 |
+
"completions/mean_length": 251.2125,
|
| 157 |
+
"completions/mean_terminated_length": 56.1,
|
| 158 |
+
"completions/min_length": 196.8,
|
| 159 |
+
"completions/min_terminated_length": 43.2,
|
| 160 |
+
"entropy": 0.8586828224360943,
|
| 161 |
+
"epoch": 0.6,
|
| 162 |
+
"frac_reward_zero_std": 0.6,
|
| 163 |
+
"grad_norm": 0.059814453125,
|
| 164 |
+
"learning_rate": 4.71e-05,
|
| 165 |
+
"loss": -0.005647056177258492,
|
| 166 |
+
"num_tokens": 208250.0,
|
| 167 |
+
"reward": 0.0625,
|
| 168 |
+
"reward_std": 0.18662600517272948,
|
| 169 |
+
"rewards/dbre_reward/mean": 0.0625,
|
| 170 |
+
"rewards/dbre_reward/std": 0.18662601709365845,
|
| 171 |
+
"step": 30,
|
| 172 |
+
"step_time": 37.554382849001556
|
| 173 |
+
},
|
| 174 |
+
{
|
| 175 |
+
"clip_ratio/high_max": 0.0,
|
| 176 |
+
"clip_ratio/high_mean": 0.0,
|
| 177 |
+
"clip_ratio/low_mean": 0.0,
|
| 178 |
+
"clip_ratio/low_min": 0.0,
|
| 179 |
+
"clip_ratio/region_mean": 0.0,
|
| 180 |
+
"completions/clipped_ratio": 0.9625,
|
| 181 |
+
"completions/max_length": 256.0,
|
| 182 |
+
"completions/max_terminated_length": 115.6,
|
| 183 |
+
"completions/mean_length": 253.625,
|
| 184 |
+
"completions/mean_terminated_length": 115.6,
|
| 185 |
+
"completions/min_length": 218.0,
|
| 186 |
+
"completions/min_terminated_length": 115.6,
|
| 187 |
+
"entropy": 0.8725819021463395,
|
| 188 |
+
"epoch": 0.7,
|
| 189 |
+
"frac_reward_zero_std": 0.3,
|
| 190 |
+
"grad_norm": 0.06787109375,
|
| 191 |
+
"learning_rate": 4.660000000000001e-05,
|
| 192 |
+
"loss": 0.0011730872094631196,
|
| 193 |
+
"num_tokens": 243020.0,
|
| 194 |
+
"reward": 0.11033750027418136,
|
| 195 |
+
"reward_std": 0.31098498702049254,
|
| 196 |
+
"rewards/dbre_reward/mean": 0.11033750027418136,
|
| 197 |
+
"rewards/dbre_reward/std": 0.31098498702049254,
|
| 198 |
+
"step": 35,
|
| 199 |
+
"step_time": 35.878880111602484
|
| 200 |
+
},
|
| 201 |
+
{
|
| 202 |
+
"clip_ratio/high_max": 0.0,
|
| 203 |
+
"clip_ratio/high_mean": 0.0,
|
| 204 |
+
"clip_ratio/low_mean": 0.0,
|
| 205 |
+
"clip_ratio/low_min": 0.0,
|
| 206 |
+
"clip_ratio/region_mean": 0.0,
|
| 207 |
+
"completions/clipped_ratio": 1.0,
|
| 208 |
+
"completions/max_length": 256.0,
|
| 209 |
+
"completions/max_terminated_length": 0.0,
|
| 210 |
+
"completions/mean_length": 256.0,
|
| 211 |
+
"completions/mean_terminated_length": 0.0,
|
| 212 |
+
"completions/min_length": 256.0,
|
| 213 |
+
"completions/min_terminated_length": 0.0,
|
| 214 |
+
"entropy": 0.9090480573475361,
|
| 215 |
+
"epoch": 0.8,
|
| 216 |
+
"frac_reward_zero_std": 0.4,
|
| 217 |
+
"grad_norm": 0.046875,
|
| 218 |
+
"learning_rate": 4.61e-05,
|
| 219 |
+
"loss": -8.940696716308593e-09,
|
| 220 |
+
"num_tokens": 277980.0,
|
| 221 |
+
"reward": 0.09995000064373016,
|
| 222 |
+
"reward_std": 0.30480254292488096,
|
| 223 |
+
"rewards/dbre_reward/mean": 0.09995000064373016,
|
| 224 |
+
"rewards/dbre_reward/std": 0.3048025548458099,
|
| 225 |
+
"step": 40,
|
| 226 |
+
"step_time": 34.05516860120406
|
| 227 |
+
},
|
| 228 |
+
{
|
| 229 |
+
"clip_ratio/high_max": 0.0,
|
| 230 |
+
"clip_ratio/high_mean": 0.0,
|
| 231 |
+
"clip_ratio/low_mean": 0.0,
|
| 232 |
+
"clip_ratio/low_min": 0.0,
|
| 233 |
+
"clip_ratio/region_mean": 0.0,
|
| 234 |
+
"completions/clipped_ratio": 0.9875,
|
| 235 |
+
"completions/max_length": 256.0,
|
| 236 |
+
"completions/max_terminated_length": 26.0,
|
| 237 |
+
"completions/mean_length": 254.425,
|
| 238 |
+
"completions/mean_terminated_length": 26.0,
|
| 239 |
+
"completions/min_length": 230.8,
|
| 240 |
+
"completions/min_terminated_length": 26.0,
|
| 241 |
+
"entropy": 0.8977848328649998,
|
| 242 |
+
"epoch": 0.9,
|
| 243 |
+
"frac_reward_zero_std": 0.6,
|
| 244 |
+
"grad_norm": 0.0,
|
| 245 |
+
"learning_rate": 4.5600000000000004e-05,
|
| 246 |
+
"loss": 1.4901161193847657e-09,
|
| 247 |
+
"num_tokens": 312814.0,
|
| 248 |
+
"reward": 0.07415000200271607,
|
| 249 |
+
"reward_std": 0.197099506855011,
|
| 250 |
+
"rewards/dbre_reward/mean": 0.07415000200271607,
|
| 251 |
+
"rewards/dbre_reward/std": 0.197099506855011,
|
| 252 |
+
"step": 45,
|
| 253 |
+
"step_time": 34.819741847200206
|
| 254 |
+
},
|
| 255 |
+
{
|
| 256 |
+
"clip_ratio/high_max": 0.0,
|
| 257 |
+
"clip_ratio/high_mean": 0.0,
|
| 258 |
+
"clip_ratio/low_mean": 0.0,
|
| 259 |
+
"clip_ratio/low_min": 0.0,
|
| 260 |
+
"clip_ratio/region_mean": 0.0,
|
| 261 |
+
"completions/clipped_ratio": 0.975,
|
| 262 |
+
"completions/max_length": 256.0,
|
| 263 |
+
"completions/max_terminated_length": 20.4,
|
| 264 |
+
"completions/mean_length": 250.875,
|
| 265 |
+
"completions/mean_terminated_length": 20.4,
|
| 266 |
+
"completions/min_length": 174.0,
|
| 267 |
+
"completions/min_terminated_length": 20.4,
|
| 268 |
+
"entropy": 1.0103468239307403,
|
| 269 |
+
"epoch": 1.0,
|
| 270 |
+
"frac_reward_zero_std": 0.8,
|
| 271 |
+
"grad_norm": 0.0,
|
| 272 |
+
"learning_rate": 4.5100000000000005e-05,
|
| 273 |
+
"loss": -2.2351741790771484e-09,
|
| 274 |
+
"num_tokens": 347364.0,
|
| 275 |
+
"reward": 0.04707500040531158,
|
| 276 |
+
"reward_std": 0.12459058165550232,
|
| 277 |
+
"rewards/dbre_reward/mean": 0.04707500040531158,
|
| 278 |
+
"rewards/dbre_reward/std": 0.12459058165550232,
|
| 279 |
+
"step": 50,
|
| 280 |
+
"step_time": 36.14326306019211
|
| 281 |
+
},
|
| 282 |
+
{
|
| 283 |
+
"clip_ratio/high_max": 0.0,
|
| 284 |
+
"clip_ratio/high_mean": 0.0,
|
| 285 |
+
"clip_ratio/low_mean": 0.0,
|
| 286 |
+
"clip_ratio/low_min": 0.0,
|
| 287 |
+
"clip_ratio/region_mean": 0.0,
|
| 288 |
+
"completions/clipped_ratio": 0.9375,
|
| 289 |
+
"completions/max_length": 256.0,
|
| 290 |
+
"completions/max_terminated_length": 41.8,
|
| 291 |
+
"completions/mean_length": 244.5625,
|
| 292 |
+
"completions/mean_terminated_length": 18.55,
|
| 293 |
+
"completions/min_length": 161.4,
|
| 294 |
+
"completions/min_terminated_length": 7.8,
|
| 295 |
+
"entropy": 0.9930311724543571,
|
| 296 |
+
"epoch": 1.1,
|
| 297 |
+
"frac_reward_zero_std": 0.6,
|
| 298 |
+
"grad_norm": 0.06396484375,
|
| 299 |
+
"learning_rate": 4.46e-05,
|
| 300 |
+
"loss": -0.017940016090869905,
|
| 301 |
+
"num_tokens": 381409.0,
|
| 302 |
+
"reward": 0.06066250056028366,
|
| 303 |
+
"reward_std": 0.2124839812517166,
|
| 304 |
+
"rewards/dbre_reward/mean": 0.06066250056028366,
|
| 305 |
+
"rewards/dbre_reward/std": 0.21248398423194886,
|
| 306 |
+
"step": 55,
|
| 307 |
+
"step_time": 28.94829335878603
|
| 308 |
+
},
|
| 309 |
+
{
|
| 310 |
+
"clip_ratio/high_max": 0.0,
|
| 311 |
+
"clip_ratio/high_mean": 0.0,
|
| 312 |
+
"clip_ratio/low_mean": 0.0,
|
| 313 |
+
"clip_ratio/low_min": 0.0,
|
| 314 |
+
"clip_ratio/region_mean": 0.0,
|
| 315 |
+
"completions/clipped_ratio": 0.9875,
|
| 316 |
+
"completions/max_length": 256.0,
|
| 317 |
+
"completions/max_terminated_length": 9.6,
|
| 318 |
+
"completions/mean_length": 253.4,
|
| 319 |
+
"completions/mean_terminated_length": 9.6,
|
| 320 |
+
"completions/min_length": 214.4,
|
| 321 |
+
"completions/min_terminated_length": 9.6,
|
| 322 |
+
"entropy": 0.8150956228375434,
|
| 323 |
+
"epoch": 1.2,
|
| 324 |
+
"frac_reward_zero_std": 0.5,
|
| 325 |
+
"grad_norm": 0.06103515625,
|
| 326 |
+
"learning_rate": 4.41e-05,
|
| 327 |
+
"loss": -4.470348358154297e-09,
|
| 328 |
+
"num_tokens": 416161.0,
|
| 329 |
+
"reward": 0.11134999990463257,
|
| 330 |
+
"reward_std": 0.2740528523921967,
|
| 331 |
+
"rewards/dbre_reward/mean": 0.11134999990463257,
|
| 332 |
+
"rewards/dbre_reward/std": 0.27405285835266113,
|
| 333 |
+
"step": 60,
|
| 334 |
+
"step_time": 27.705245727201692
|
| 335 |
+
},
|
| 336 |
+
{
|
| 337 |
+
"clip_ratio/high_max": 0.0,
|
| 338 |
+
"clip_ratio/high_mean": 0.0,
|
| 339 |
+
"clip_ratio/low_mean": 0.0,
|
| 340 |
+
"clip_ratio/low_min": 0.0,
|
| 341 |
+
"clip_ratio/region_mean": 0.0,
|
| 342 |
+
"completions/clipped_ratio": 0.975,
|
| 343 |
+
"completions/max_length": 256.0,
|
| 344 |
+
"completions/max_terminated_length": 41.0,
|
| 345 |
+
"completions/mean_length": 252.1625,
|
| 346 |
+
"completions/mean_terminated_length": 41.0,
|
| 347 |
+
"completions/min_length": 194.6,
|
| 348 |
+
"completions/min_terminated_length": 41.0,
|
| 349 |
+
"entropy": 0.9309644259512424,
|
| 350 |
+
"epoch": 1.3,
|
| 351 |
+
"frac_reward_zero_std": 0.7,
|
| 352 |
+
"grad_norm": 0.0,
|
| 353 |
+
"learning_rate": 4.36e-05,
|
| 354 |
+
"loss": -8.940696716308593e-09,
|
| 355 |
+
"num_tokens": 450814.0,
|
| 356 |
+
"reward": 0.0875,
|
| 357 |
+
"reward_std": 0.21124515533447266,
|
| 358 |
+
"rewards/dbre_reward/mean": 0.0875,
|
| 359 |
+
"rewards/dbre_reward/std": 0.21124515533447266,
|
| 360 |
+
"step": 65,
|
| 361 |
+
"step_time": 27.76657635839365
|
| 362 |
+
},
|
| 363 |
+
{
|
| 364 |
+
"clip_ratio/high_max": 0.0,
|
| 365 |
+
"clip_ratio/high_mean": 0.0,
|
| 366 |
+
"clip_ratio/low_mean": 0.0,
|
| 367 |
+
"clip_ratio/low_min": 0.0,
|
| 368 |
+
"clip_ratio/region_mean": 0.0,
|
| 369 |
+
"completions/clipped_ratio": 0.9875,
|
| 370 |
+
"completions/max_length": 256.0,
|
| 371 |
+
"completions/max_terminated_length": 30.4,
|
| 372 |
+
"completions/mean_length": 254.7,
|
| 373 |
+
"completions/mean_terminated_length": 30.4,
|
| 374 |
+
"completions/min_length": 235.2,
|
| 375 |
+
"completions/min_terminated_length": 30.4,
|
| 376 |
+
"entropy": 0.8473479233682155,
|
| 377 |
+
"epoch": 1.4,
|
| 378 |
+
"frac_reward_zero_std": 0.6,
|
| 379 |
+
"grad_norm": 0.05078125,
|
| 380 |
+
"learning_rate": 4.3100000000000004e-05,
|
| 381 |
+
"loss": -2.2351741790771484e-09,
|
| 382 |
+
"num_tokens": 485670.0,
|
| 383 |
+
"reward": 0.07392499968409538,
|
| 384 |
+
"reward_std": 0.1946355789899826,
|
| 385 |
+
"rewards/dbre_reward/mean": 0.07392499968409538,
|
| 386 |
+
"rewards/dbre_reward/std": 0.1946355879306793,
|
| 387 |
+
"step": 70,
|
| 388 |
+
"step_time": 27.701584570185513
|
| 389 |
+
},
|
| 390 |
+
{
|
| 391 |
+
"clip_ratio/high_max": 0.0,
|
| 392 |
+
"clip_ratio/high_mean": 0.0,
|
| 393 |
+
"clip_ratio/low_mean": 0.0,
|
| 394 |
+
"clip_ratio/low_min": 0.0,
|
| 395 |
+
"clip_ratio/region_mean": 0.0,
|
| 396 |
+
"completions/clipped_ratio": 0.975,
|
| 397 |
+
"completions/max_length": 256.0,
|
| 398 |
+
"completions/max_terminated_length": 51.0,
|
| 399 |
+
"completions/mean_length": 252.7875,
|
| 400 |
+
"completions/mean_terminated_length": 51.0,
|
| 401 |
+
"completions/min_length": 204.6,
|
| 402 |
+
"completions/min_terminated_length": 51.0,
|
| 403 |
+
"entropy": 0.8822006396949291,
|
| 404 |
+
"epoch": 1.5,
|
| 405 |
+
"frac_reward_zero_std": 0.4,
|
| 406 |
+
"grad_norm": 0.046875,
|
| 407 |
+
"learning_rate": 4.26e-05,
|
| 408 |
+
"loss": -0.004587128758430481,
|
| 409 |
+
"num_tokens": 520373.0,
|
| 410 |
+
"reward": 0.09866249859333039,
|
| 411 |
+
"reward_std": 0.29479086995124815,
|
| 412 |
+
"rewards/dbre_reward/mean": 0.09866249859333039,
|
| 413 |
+
"rewards/dbre_reward/std": 0.29479087591171266,
|
| 414 |
+
"step": 75,
|
| 415 |
+
"step_time": 27.723117466596886
|
| 416 |
+
},
|
| 417 |
+
{
|
| 418 |
+
"clip_ratio/high_max": 0.0,
|
| 419 |
+
"clip_ratio/high_mean": 0.0,
|
| 420 |
+
"clip_ratio/low_mean": 0.0,
|
| 421 |
+
"clip_ratio/low_min": 0.0,
|
| 422 |
+
"clip_ratio/region_mean": 0.0,
|
| 423 |
+
"completions/clipped_ratio": 1.0,
|
| 424 |
+
"completions/max_length": 256.0,
|
| 425 |
+
"completions/max_terminated_length": 0.0,
|
| 426 |
+
"completions/mean_length": 256.0,
|
| 427 |
+
"completions/mean_terminated_length": 0.0,
|
| 428 |
+
"completions/min_length": 256.0,
|
| 429 |
+
"completions/min_terminated_length": 0.0,
|
| 430 |
+
"entropy": 0.8492010429501533,
|
| 431 |
+
"epoch": 1.6,
|
| 432 |
+
"frac_reward_zero_std": 0.4,
|
| 433 |
+
"grad_norm": 0.052734375,
|
| 434 |
+
"learning_rate": 4.21e-05,
|
| 435 |
+
"loss": -4.470348358154297e-09,
|
| 436 |
+
"num_tokens": 555333.0,
|
| 437 |
+
"reward": 0.13631249964237213,
|
| 438 |
+
"reward_std": 0.3453687012195587,
|
| 439 |
+
"rewards/dbre_reward/mean": 0.13631249964237213,
|
| 440 |
+
"rewards/dbre_reward/std": 0.34536872506141664,
|
| 441 |
+
"step": 80,
|
| 442 |
+
"step_time": 27.74355768300593
|
| 443 |
+
},
|
| 444 |
+
{
|
| 445 |
+
"clip_ratio/high_max": 0.0,
|
| 446 |
+
"clip_ratio/high_mean": 0.0,
|
| 447 |
+
"clip_ratio/low_mean": 0.0,
|
| 448 |
+
"clip_ratio/low_min": 0.0,
|
| 449 |
+
"clip_ratio/region_mean": 0.0,
|
| 450 |
+
"completions/clipped_ratio": 0.9875,
|
| 451 |
+
"completions/max_length": 256.0,
|
| 452 |
+
"completions/max_terminated_length": 23.0,
|
| 453 |
+
"completions/mean_length": 254.2375,
|
| 454 |
+
"completions/mean_terminated_length": 23.0,
|
| 455 |
+
"completions/min_length": 227.8,
|
| 456 |
+
"completions/min_terminated_length": 23.0,
|
| 457 |
+
"entropy": 0.9008926346898078,
|
| 458 |
+
"epoch": 1.7,
|
| 459 |
+
"frac_reward_zero_std": 0.4,
|
| 460 |
+
"grad_norm": 0.053466796875,
|
| 461 |
+
"learning_rate": 4.16e-05,
|
| 462 |
+
"loss": -0.0038484178483486177,
|
| 463 |
+
"num_tokens": 590152.0,
|
| 464 |
+
"reward": 0.13577499985694885,
|
| 465 |
+
"reward_std": 0.3495619535446167,
|
| 466 |
+
"rewards/dbre_reward/mean": 0.13577499985694885,
|
| 467 |
+
"rewards/dbre_reward/std": 0.34956197142601014,
|
| 468 |
+
"step": 85,
|
| 469 |
+
"step_time": 27.809819040997535
|
| 470 |
+
},
|
| 471 |
+
{
|
| 472 |
+
"clip_ratio/high_max": 0.0,
|
| 473 |
+
"clip_ratio/high_mean": 0.0,
|
| 474 |
+
"clip_ratio/low_mean": 0.0,
|
| 475 |
+
"clip_ratio/low_min": 0.0,
|
| 476 |
+
"clip_ratio/region_mean": 0.0,
|
| 477 |
+
"completions/clipped_ratio": 0.975,
|
| 478 |
+
"completions/max_length": 256.0,
|
| 479 |
+
"completions/max_terminated_length": 34.4,
|
| 480 |
+
"completions/mean_length": 253.7375,
|
| 481 |
+
"completions/mean_terminated_length": 33.1,
|
| 482 |
+
"completions/min_length": 236.6,
|
| 483 |
+
"completions/min_terminated_length": 31.8,
|
| 484 |
+
"entropy": 0.975049901008606,
|
| 485 |
+
"epoch": 1.8,
|
| 486 |
+
"frac_reward_zero_std": 0.3,
|
| 487 |
+
"grad_norm": 0.08447265625,
|
| 488 |
+
"learning_rate": 4.11e-05,
|
| 489 |
+
"loss": -0.0023170128464698792,
|
| 490 |
+
"num_tokens": 624931.0,
|
| 491 |
+
"reward": 0.14498749673366546,
|
| 492 |
+
"reward_std": 0.3355918139219284,
|
| 493 |
+
"rewards/dbre_reward/mean": 0.14498749673366546,
|
| 494 |
+
"rewards/dbre_reward/std": 0.3355918198823929,
|
| 495 |
+
"step": 90,
|
| 496 |
+
"step_time": 27.872812610585243
|
| 497 |
+
},
|
| 498 |
+
{
|
| 499 |
+
"clip_ratio/high_max": 0.0,
|
| 500 |
+
"clip_ratio/high_mean": 0.0,
|
| 501 |
+
"clip_ratio/low_mean": 0.0,
|
| 502 |
+
"clip_ratio/low_min": 0.0,
|
| 503 |
+
"clip_ratio/region_mean": 0.0,
|
| 504 |
+
"completions/clipped_ratio": 0.95,
|
| 505 |
+
"completions/max_length": 256.0,
|
| 506 |
+
"completions/max_terminated_length": 75.6,
|
| 507 |
+
"completions/mean_length": 250.775,
|
| 508 |
+
"completions/mean_terminated_length": 60.6,
|
| 509 |
+
"completions/min_length": 199.2,
|
| 510 |
+
"completions/min_terminated_length": 45.6,
|
| 511 |
+
"entropy": 0.9378853186964988,
|
| 512 |
+
"epoch": 1.9,
|
| 513 |
+
"frac_reward_zero_std": 0.4,
|
| 514 |
+
"grad_norm": 0.0576171875,
|
| 515 |
+
"learning_rate": 4.0600000000000004e-05,
|
| 516 |
+
"loss": -0.010608357191085816,
|
| 517 |
+
"num_tokens": 659473.0,
|
| 518 |
+
"reward": 0.13513749986886978,
|
| 519 |
+
"reward_std": 0.3332351267337799,
|
| 520 |
+
"rewards/dbre_reward/mean": 0.13513749986886978,
|
| 521 |
+
"rewards/dbre_reward/std": 0.3332351326942444,
|
| 522 |
+
"step": 95,
|
| 523 |
+
"step_time": 27.71230500699894
|
| 524 |
+
},
|
| 525 |
+
{
|
| 526 |
+
"clip_ratio/high_max": 0.0,
|
| 527 |
+
"clip_ratio/high_mean": 0.0,
|
| 528 |
+
"clip_ratio/low_mean": 0.0,
|
| 529 |
+
"clip_ratio/low_min": 0.0,
|
| 530 |
+
"clip_ratio/region_mean": 0.0,
|
| 531 |
+
"completions/clipped_ratio": 0.975,
|
| 532 |
+
"completions/max_length": 256.0,
|
| 533 |
+
"completions/max_terminated_length": 80.0,
|
| 534 |
+
"completions/mean_length": 254.6,
|
| 535 |
+
"completions/mean_terminated_length": 80.0,
|
| 536 |
+
"completions/min_length": 233.6,
|
| 537 |
+
"completions/min_terminated_length": 80.0,
|
| 538 |
+
"entropy": 0.8716862492263318,
|
| 539 |
+
"epoch": 2.0,
|
| 540 |
+
"frac_reward_zero_std": 0.4,
|
| 541 |
+
"grad_norm": 0.06396484375,
|
| 542 |
+
"learning_rate": 4.0100000000000006e-05,
|
| 543 |
+
"loss": 0.0028070926666259764,
|
| 544 |
+
"num_tokens": 694321.0,
|
| 545 |
+
"reward": 0.13687500059604646,
|
| 546 |
+
"reward_std": 0.33139119744300843,
|
| 547 |
+
"rewards/dbre_reward/mean": 0.13687500059604646,
|
| 548 |
+
"rewards/dbre_reward/std": 0.3313912093639374,
|
| 549 |
+
"step": 100,
|
| 550 |
+
"step_time": 27.905723336405934
|
| 551 |
+
},
|
| 552 |
+
{
|
| 553 |
+
"clip_ratio/high_max": 0.0,
|
| 554 |
+
"clip_ratio/high_mean": 0.0,
|
| 555 |
+
"clip_ratio/low_mean": 0.0,
|
| 556 |
+
"clip_ratio/low_min": 0.0,
|
| 557 |
+
"clip_ratio/region_mean": 0.0,
|
| 558 |
+
"completions/clipped_ratio": 0.95,
|
| 559 |
+
"completions/max_length": 256.0,
|
| 560 |
+
"completions/max_terminated_length": 138.6,
|
| 561 |
+
"completions/mean_length": 251.8625,
|
| 562 |
+
"completions/mean_terminated_length": 138.6,
|
| 563 |
+
"completions/min_length": 189.8,
|
| 564 |
+
"completions/min_terminated_length": 138.6,
|
| 565 |
+
"entropy": 0.9404824480414391,
|
| 566 |
+
"epoch": 2.1,
|
| 567 |
+
"frac_reward_zero_std": 0.5,
|
| 568 |
+
"grad_norm": 0.060791015625,
|
| 569 |
+
"learning_rate": 3.960000000000001e-05,
|
| 570 |
+
"loss": -0.004697377979755402,
|
| 571 |
+
"num_tokens": 728950.0,
|
| 572 |
+
"reward": 0.08519999980926514,
|
| 573 |
+
"reward_std": 0.24869290590286255,
|
| 574 |
+
"rewards/dbre_reward/mean": 0.08519999980926514,
|
| 575 |
+
"rewards/dbre_reward/std": 0.24869290590286255,
|
| 576 |
+
"step": 105,
|
| 577 |
+
"step_time": 27.779890004795742
|
| 578 |
+
},
|
| 579 |
+
{
|
| 580 |
+
"clip_ratio/high_max": 0.0,
|
| 581 |
+
"clip_ratio/high_mean": 0.0,
|
| 582 |
+
"clip_ratio/low_mean": 0.0,
|
| 583 |
+
"clip_ratio/low_min": 0.0,
|
| 584 |
+
"clip_ratio/region_mean": 0.0,
|
| 585 |
+
"completions/clipped_ratio": 0.9625,
|
| 586 |
+
"completions/max_length": 256.0,
|
| 587 |
+
"completions/max_terminated_length": 21.8,
|
| 588 |
+
"completions/mean_length": 248.425,
|
| 589 |
+
"completions/mean_terminated_length": 19.9,
|
| 590 |
+
"completions/min_length": 171.6,
|
| 591 |
+
"completions/min_terminated_length": 18.0,
|
| 592 |
+
"entropy": 1.0435175843536855,
|
| 593 |
+
"epoch": 2.2,
|
| 594 |
+
"frac_reward_zero_std": 0.3,
|
| 595 |
+
"grad_norm": 0.134765625,
|
| 596 |
+
"learning_rate": 3.91e-05,
|
| 597 |
+
"loss": -0.007500007748603821,
|
| 598 |
+
"num_tokens": 763304.0,
|
| 599 |
+
"reward": 0.13616250157356263,
|
| 600 |
+
"reward_std": 0.3243652701377869,
|
| 601 |
+
"rewards/dbre_reward/mean": 0.13616250157356263,
|
| 602 |
+
"rewards/dbre_reward/std": 0.3243652701377869,
|
| 603 |
+
"step": 110,
|
| 604 |
+
"step_time": 27.559503265400416
|
| 605 |
+
},
|
| 606 |
+
{
|
| 607 |
+
"clip_ratio/high_max": 0.0,
|
| 608 |
+
"clip_ratio/high_mean": 0.0,
|
| 609 |
+
"clip_ratio/low_mean": 0.0,
|
| 610 |
+
"clip_ratio/low_min": 0.0,
|
| 611 |
+
"clip_ratio/region_mean": 0.0,
|
| 612 |
+
"completions/clipped_ratio": 0.9875,
|
| 613 |
+
"completions/max_length": 256.0,
|
| 614 |
+
"completions/max_terminated_length": 10.0,
|
| 615 |
+
"completions/mean_length": 253.425,
|
| 616 |
+
"completions/mean_terminated_length": 10.0,
|
| 617 |
+
"completions/min_length": 214.8,
|
| 618 |
+
"completions/min_terminated_length": 10.0,
|
| 619 |
+
"entropy": 0.9914286866784096,
|
| 620 |
+
"epoch": 2.3,
|
| 621 |
+
"frac_reward_zero_std": 0.2,
|
| 622 |
+
"grad_norm": 0.0830078125,
|
| 623 |
+
"learning_rate": 3.86e-05,
|
| 624 |
+
"loss": -0.0037435129284858703,
|
| 625 |
+
"num_tokens": 798058.0,
|
| 626 |
+
"reward": 0.17202500104904175,
|
| 627 |
+
"reward_std": 0.37993268966674804,
|
| 628 |
+
"rewards/dbre_reward/mean": 0.17202500104904175,
|
| 629 |
+
"rewards/dbre_reward/std": 0.37993271350860597,
|
| 630 |
+
"step": 115,
|
| 631 |
+
"step_time": 27.579605548398103
|
| 632 |
+
},
|
| 633 |
+
{
|
| 634 |
+
"clip_ratio/high_max": 0.0,
|
| 635 |
+
"clip_ratio/high_mean": 0.0,
|
| 636 |
+
"clip_ratio/low_mean": 0.0,
|
| 637 |
+
"clip_ratio/low_min": 0.0,
|
| 638 |
+
"clip_ratio/region_mean": 0.0,
|
| 639 |
+
"completions/clipped_ratio": 0.95,
|
| 640 |
+
"completions/max_length": 256.0,
|
| 641 |
+
"completions/max_terminated_length": 128.6,
|
| 642 |
+
"completions/mean_length": 252.5,
|
| 643 |
+
"completions/mean_terminated_length": 117.6,
|
| 644 |
+
"completions/min_length": 209.0,
|
| 645 |
+
"completions/min_terminated_length": 106.6,
|
| 646 |
+
"entropy": 1.0246646717190742,
|
| 647 |
+
"epoch": 2.4,
|
| 648 |
+
"frac_reward_zero_std": 0.2,
|
| 649 |
+
"grad_norm": 0.0927734375,
|
| 650 |
+
"learning_rate": 3.8100000000000005e-05,
|
| 651 |
+
"loss": -0.004933140054345131,
|
| 652 |
+
"num_tokens": 832738.0,
|
| 653 |
+
"reward": 0.18355000019073486,
|
| 654 |
+
"reward_std": 0.36511489748954773,
|
| 655 |
+
"rewards/dbre_reward/mean": 0.18355000019073486,
|
| 656 |
+
"rewards/dbre_reward/std": 0.36511489748954773,
|
| 657 |
+
"step": 120,
|
| 658 |
+
"step_time": 27.58797151680337
|
| 659 |
+
},
|
| 660 |
+
{
|
| 661 |
+
"clip_ratio/high_max": 0.0,
|
| 662 |
+
"clip_ratio/high_mean": 0.0,
|
| 663 |
+
"clip_ratio/low_mean": 0.0,
|
| 664 |
+
"clip_ratio/low_min": 0.0,
|
| 665 |
+
"clip_ratio/region_mean": 0.0,
|
| 666 |
+
"completions/clipped_ratio": 0.9875,
|
| 667 |
+
"completions/max_length": 256.0,
|
| 668 |
+
"completions/max_terminated_length": 5.2,
|
| 669 |
+
"completions/mean_length": 253.125,
|
| 670 |
+
"completions/mean_terminated_length": 5.2,
|
| 671 |
+
"completions/min_length": 210.0,
|
| 672 |
+
"completions/min_terminated_length": 5.2,
|
| 673 |
+
"entropy": 0.9524292007088662,
|
| 674 |
+
"epoch": 2.5,
|
| 675 |
+
"frac_reward_zero_std": 0.3,
|
| 676 |
+
"grad_norm": 0.083984375,
|
| 677 |
+
"learning_rate": 3.76e-05,
|
| 678 |
+
"loss": -0.004205547273159027,
|
| 679 |
+
"num_tokens": 867468.0,
|
| 680 |
+
"reward": 0.22202499732375144,
|
| 681 |
+
"reward_std": 0.3661981552839279,
|
| 682 |
+
"rewards/dbre_reward/mean": 0.22202499732375144,
|
| 683 |
+
"rewards/dbre_reward/std": 0.3661981612443924,
|
| 684 |
+
"step": 125,
|
| 685 |
+
"step_time": 27.638399406394456
|
| 686 |
+
},
|
| 687 |
+
{
|
| 688 |
+
"clip_ratio/high_max": 0.0,
|
| 689 |
+
"clip_ratio/high_mean": 0.0,
|
| 690 |
+
"clip_ratio/low_mean": 0.0,
|
| 691 |
+
"clip_ratio/low_min": 0.0,
|
| 692 |
+
"clip_ratio/region_mean": 0.0,
|
| 693 |
+
"completions/clipped_ratio": 0.9875,
|
| 694 |
+
"completions/max_length": 256.0,
|
| 695 |
+
"completions/max_terminated_length": 34.4,
|
| 696 |
+
"completions/mean_length": 254.95,
|
| 697 |
+
"completions/mean_terminated_length": 34.4,
|
| 698 |
+
"completions/min_length": 239.2,
|
| 699 |
+
"completions/min_terminated_length": 34.4,
|
| 700 |
+
"entropy": 0.9268762037158013,
|
| 701 |
+
"epoch": 2.6,
|
| 702 |
+
"frac_reward_zero_std": 0.4,
|
| 703 |
+
"grad_norm": 0.060546875,
|
| 704 |
+
"learning_rate": 3.71e-05,
|
| 705 |
+
"loss": -0.002260996401309967,
|
| 706 |
+
"num_tokens": 902344.0,
|
| 707 |
+
"reward": 0.19577499628067016,
|
| 708 |
+
"reward_std": 0.3975414574146271,
|
| 709 |
+
"rewards/dbre_reward/mean": 0.19577499628067016,
|
| 710 |
+
"rewards/dbre_reward/std": 0.397541481256485,
|
| 711 |
+
"step": 130,
|
| 712 |
+
"step_time": 27.68072868139425
|
| 713 |
+
},
|
| 714 |
+
{
|
| 715 |
+
"clip_ratio/high_max": 0.0,
|
| 716 |
+
"clip_ratio/high_mean": 0.0,
|
| 717 |
+
"clip_ratio/low_mean": 0.0,
|
| 718 |
+
"clip_ratio/low_min": 0.0,
|
| 719 |
+
"clip_ratio/region_mean": 0.0,
|
| 720 |
+
"completions/clipped_ratio": 0.9875,
|
| 721 |
+
"completions/max_length": 256.0,
|
| 722 |
+
"completions/max_terminated_length": 26.6,
|
| 723 |
+
"completions/mean_length": 254.4625,
|
| 724 |
+
"completions/mean_terminated_length": 26.6,
|
| 725 |
+
"completions/min_length": 231.4,
|
| 726 |
+
"completions/min_terminated_length": 26.6,
|
| 727 |
+
"entropy": 0.9508431695401669,
|
| 728 |
+
"epoch": 2.7,
|
| 729 |
+
"frac_reward_zero_std": 0.1,
|
| 730 |
+
"grad_norm": 0.1025390625,
|
| 731 |
+
"learning_rate": 3.66e-05,
|
| 732 |
+
"loss": -0.004485464096069336,
|
| 733 |
+
"num_tokens": 937181.0,
|
| 734 |
+
"reward": 0.1941875010728836,
|
| 735 |
+
"reward_std": 0.39680722951889036,
|
| 736 |
+
"rewards/dbre_reward/mean": 0.1941875010728836,
|
| 737 |
+
"rewards/dbre_reward/std": 0.39680724740028384,
|
| 738 |
+
"step": 135,
|
| 739 |
+
"step_time": 27.79836461079831
|
| 740 |
+
},
|
| 741 |
+
{
|
| 742 |
+
"clip_ratio/high_max": 0.0,
|
| 743 |
+
"clip_ratio/high_mean": 0.0,
|
| 744 |
+
"clip_ratio/low_mean": 0.0,
|
| 745 |
+
"clip_ratio/low_min": 0.0,
|
| 746 |
+
"clip_ratio/region_mean": 0.0,
|
| 747 |
+
"completions/clipped_ratio": 1.0,
|
| 748 |
+
"completions/max_length": 256.0,
|
| 749 |
+
"completions/max_terminated_length": 0.0,
|
| 750 |
+
"completions/mean_length": 256.0,
|
| 751 |
+
"completions/mean_terminated_length": 0.0,
|
| 752 |
+
"completions/min_length": 256.0,
|
| 753 |
+
"completions/min_terminated_length": 0.0,
|
| 754 |
+
"entropy": 0.9166437476873398,
|
| 755 |
+
"epoch": 2.8,
|
| 756 |
+
"frac_reward_zero_std": 0.1,
|
| 757 |
+
"grad_norm": 0.078125,
|
| 758 |
+
"learning_rate": 3.61e-05,
|
| 759 |
+
"loss": 2.9802322387695314e-09,
|
| 760 |
+
"num_tokens": 972141.0,
|
| 761 |
+
"reward": 0.1720750018954277,
|
| 762 |
+
"reward_std": 0.3809880971908569,
|
| 763 |
+
"rewards/dbre_reward/mean": 0.1720750018954277,
|
| 764 |
+
"rewards/dbre_reward/std": 0.3809881091117859,
|
| 765 |
+
"step": 140,
|
| 766 |
+
"step_time": 27.89382164280105
|
| 767 |
+
},
|
| 768 |
+
{
|
| 769 |
+
"clip_ratio/high_max": 0.0,
|
| 770 |
+
"clip_ratio/high_mean": 0.0,
|
| 771 |
+
"clip_ratio/low_mean": 0.0,
|
| 772 |
+
"clip_ratio/low_min": 0.0,
|
| 773 |
+
"clip_ratio/region_mean": 0.0,
|
| 774 |
+
"completions/clipped_ratio": 0.95,
|
| 775 |
+
"completions/max_length": 256.0,
|
| 776 |
+
"completions/max_terminated_length": 117.2,
|
| 777 |
+
"completions/mean_length": 252.3125,
|
| 778 |
+
"completions/mean_terminated_length": 108.2,
|
| 779 |
+
"completions/min_length": 201.6,
|
| 780 |
+
"completions/min_terminated_length": 99.2,
|
| 781 |
+
"entropy": 1.039070624113083,
|
| 782 |
+
"epoch": 2.9,
|
| 783 |
+
"frac_reward_zero_std": 0.2,
|
| 784 |
+
"grad_norm": 0.08740234375,
|
| 785 |
+
"learning_rate": 3.56e-05,
|
| 786 |
+
"loss": 0.003438304364681244,
|
| 787 |
+
"num_tokens": 1006806.0,
|
| 788 |
+
"reward": 0.2208999961614609,
|
| 789 |
+
"reward_std": 0.401767635345459,
|
| 790 |
+
"rewards/dbre_reward/mean": 0.2208999961614609,
|
| 791 |
+
"rewards/dbre_reward/std": 0.4017676472663879,
|
| 792 |
+
"step": 145,
|
| 793 |
+
"step_time": 27.960635445202932
|
| 794 |
+
},
|
| 795 |
+
{
|
| 796 |
+
"clip_ratio/high_max": 0.0,
|
| 797 |
+
"clip_ratio/high_mean": 0.0,
|
| 798 |
+
"clip_ratio/low_mean": 0.0,
|
| 799 |
+
"clip_ratio/low_min": 0.0,
|
| 800 |
+
"clip_ratio/region_mean": 0.0,
|
| 801 |
+
"completions/clipped_ratio": 0.975,
|
| 802 |
+
"completions/max_length": 256.0,
|
| 803 |
+
"completions/max_terminated_length": 40.2,
|
| 804 |
+
"completions/mean_length": 253.0625,
|
| 805 |
+
"completions/mean_terminated_length": 27.7,
|
| 806 |
+
"completions/min_length": 220.0,
|
| 807 |
+
"completions/min_terminated_length": 15.2,
|
| 808 |
+
"entropy": 0.9523506201803684,
|
| 809 |
+
"epoch": 3.0,
|
| 810 |
+
"frac_reward_zero_std": 0.1,
|
| 811 |
+
"grad_norm": 0.0966796875,
|
| 812 |
+
"learning_rate": 3.51e-05,
|
| 813 |
+
"loss": -0.006567706167697906,
|
| 814 |
+
"num_tokens": 1041531.0,
|
| 815 |
+
"reward": 0.254237499833107,
|
| 816 |
+
"reward_std": 0.4319828271865845,
|
| 817 |
+
"rewards/dbre_reward/mean": 0.254237499833107,
|
| 818 |
+
"rewards/dbre_reward/std": 0.4319828271865845,
|
| 819 |
+
"step": 150,
|
| 820 |
+
"step_time": 33.83420682080032
|
| 821 |
+
},
|
| 822 |
+
{
|
| 823 |
+
"clip_ratio/high_max": 0.0,
|
| 824 |
+
"clip_ratio/high_mean": 0.0,
|
| 825 |
+
"clip_ratio/low_mean": 0.0,
|
| 826 |
+
"clip_ratio/low_min": 0.0,
|
| 827 |
+
"clip_ratio/region_mean": 0.0,
|
| 828 |
+
"completions/clipped_ratio": 0.95,
|
| 829 |
+
"completions/max_length": 256.0,
|
| 830 |
+
"completions/max_terminated_length": 135.8,
|
| 831 |
+
"completions/mean_length": 254.4875,
|
| 832 |
+
"completions/mean_terminated_length": 134.9,
|
| 833 |
+
"completions/min_length": 236.4,
|
| 834 |
+
"completions/min_terminated_length": 134.0,
|
| 835 |
+
"entropy": 0.9468083322048187,
|
| 836 |
+
"epoch": 3.1,
|
| 837 |
+
"frac_reward_zero_std": 0.2,
|
| 838 |
+
"grad_norm": 0.0703125,
|
| 839 |
+
"learning_rate": 3.46e-05,
|
| 840 |
+
"loss": -0.003655475750565529,
|
| 841 |
+
"num_tokens": 1076370.0,
|
| 842 |
+
"reward": 0.2434374988079071,
|
| 843 |
+
"reward_std": 0.4292252540588379,
|
| 844 |
+
"rewards/dbre_reward/mean": 0.2434374988079071,
|
| 845 |
+
"rewards/dbre_reward/std": 0.4292252600193024,
|
| 846 |
+
"step": 155,
|
| 847 |
+
"step_time": 35.7378796559904
|
| 848 |
+
},
|
| 849 |
+
{
|
| 850 |
+
"clip_ratio/high_max": 0.0,
|
| 851 |
+
"clip_ratio/high_mean": 0.0,
|
| 852 |
+
"clip_ratio/low_mean": 0.0,
|
| 853 |
+
"clip_ratio/low_min": 0.0,
|
| 854 |
+
"clip_ratio/region_mean": 0.0,
|
| 855 |
+
"completions/clipped_ratio": 0.9875,
|
| 856 |
+
"completions/max_length": 256.0,
|
| 857 |
+
"completions/max_terminated_length": 10.4,
|
| 858 |
+
"completions/mean_length": 253.45,
|
| 859 |
+
"completions/mean_terminated_length": 10.4,
|
| 860 |
+
"completions/min_length": 215.2,
|
| 861 |
+
"completions/min_terminated_length": 10.4,
|
| 862 |
+
"entropy": 0.9374438695609569,
|
| 863 |
+
"epoch": 3.2,
|
| 864 |
+
"frac_reward_zero_std": 0.3,
|
| 865 |
+
"grad_norm": 0.0693359375,
|
| 866 |
+
"learning_rate": 3.41e-05,
|
| 867 |
+
"loss": -0.005657447874546051,
|
| 868 |
+
"num_tokens": 1111126.0,
|
| 869 |
+
"reward": 0.2539374977350235,
|
| 870 |
+
"reward_std": 0.4268993496894836,
|
| 871 |
+
"rewards/dbre_reward/mean": 0.2539374977350235,
|
| 872 |
+
"rewards/dbre_reward/std": 0.4268993675708771,
|
| 873 |
+
"step": 160,
|
| 874 |
+
"step_time": 33.11068361419602
|
| 875 |
+
},
|
| 876 |
+
{
|
| 877 |
+
"clip_ratio/high_max": 0.0,
|
| 878 |
+
"clip_ratio/high_mean": 0.0,
|
| 879 |
+
"clip_ratio/low_mean": 0.0,
|
| 880 |
+
"clip_ratio/low_min": 0.0,
|
| 881 |
+
"clip_ratio/region_mean": 0.0,
|
| 882 |
+
"completions/clipped_ratio": 0.925,
|
| 883 |
+
"completions/max_length": 256.0,
|
| 884 |
+
"completions/max_terminated_length": 97.2,
|
| 885 |
+
"completions/mean_length": 244.7,
|
| 886 |
+
"completions/mean_terminated_length": 95.4,
|
| 887 |
+
"completions/min_length": 93.6,
|
| 888 |
+
"completions/min_terminated_length": 93.6,
|
| 889 |
+
"entropy": 1.0759656712412835,
|
| 890 |
+
"epoch": 3.3,
|
| 891 |
+
"frac_reward_zero_std": 0.1,
|
| 892 |
+
"grad_norm": 0.09716796875,
|
| 893 |
+
"learning_rate": 3.3600000000000004e-05,
|
| 894 |
+
"loss": 0.0013453811407089233,
|
| 895 |
+
"num_tokens": 1145182.0,
|
| 896 |
+
"reward": 0.1820499964058399,
|
| 897 |
+
"reward_std": 0.37800283133983614,
|
| 898 |
+
"rewards/dbre_reward/mean": 0.1820499964058399,
|
| 899 |
+
"rewards/dbre_reward/std": 0.37800286114215853,
|
| 900 |
+
"step": 165,
|
| 901 |
+
"step_time": 33.76342459159205
|
| 902 |
+
},
|
| 903 |
+
{
|
| 904 |
+
"clip_ratio/high_max": 0.0,
|
| 905 |
+
"clip_ratio/high_mean": 0.0,
|
| 906 |
+
"clip_ratio/low_mean": 0.0,
|
| 907 |
+
"clip_ratio/low_min": 0.0,
|
| 908 |
+
"clip_ratio/region_mean": 0.0,
|
| 909 |
+
"completions/clipped_ratio": 0.925,
|
| 910 |
+
"completions/max_length": 256.0,
|
| 911 |
+
"completions/max_terminated_length": 169.2,
|
| 912 |
+
"completions/mean_length": 250.3125,
|
| 913 |
+
"completions/mean_terminated_length": 168.7,
|
| 914 |
+
"completions/min_length": 168.2,
|
| 915 |
+
"completions/min_terminated_length": 168.2,
|
| 916 |
+
"entropy": 0.9453672260046005,
|
| 917 |
+
"epoch": 3.4,
|
| 918 |
+
"frac_reward_zero_std": 0.1,
|
| 919 |
+
"grad_norm": 0.0927734375,
|
| 920 |
+
"learning_rate": 3.3100000000000005e-05,
|
| 921 |
+
"loss": -0.001752069965004921,
|
| 922 |
+
"num_tokens": 1179687.0,
|
| 923 |
+
"reward": 0.21851249933242797,
|
| 924 |
+
"reward_std": 0.4150461137294769,
|
| 925 |
+
"rewards/dbre_reward/mean": 0.21851249933242797,
|
| 926 |
+
"rewards/dbre_reward/std": 0.4150461256504059,
|
| 927 |
+
"step": 170,
|
| 928 |
+
"step_time": 35.29977663640748
|
| 929 |
+
},
|
| 930 |
+
{
|
| 931 |
+
"clip_ratio/high_max": 0.0,
|
| 932 |
+
"clip_ratio/high_mean": 0.0,
|
| 933 |
+
"clip_ratio/low_mean": 0.0,
|
| 934 |
+
"clip_ratio/low_min": 0.0,
|
| 935 |
+
"clip_ratio/region_mean": 0.0,
|
| 936 |
+
"completions/clipped_ratio": 0.9625,
|
| 937 |
+
"completions/max_length": 256.0,
|
| 938 |
+
"completions/max_terminated_length": 82.6,
|
| 939 |
+
"completions/mean_length": 251.5625,
|
| 940 |
+
"completions/mean_terminated_length": 82.6,
|
| 941 |
+
"completions/min_length": 185.0,
|
| 942 |
+
"completions/min_terminated_length": 82.6,
|
| 943 |
+
"entropy": 0.9627169869840145,
|
| 944 |
+
"epoch": 3.5,
|
| 945 |
+
"frac_reward_zero_std": 0.0,
|
| 946 |
+
"grad_norm": 0.091796875,
|
| 947 |
+
"learning_rate": 3.26e-05,
|
| 948 |
+
"loss": -0.010714849084615707,
|
| 949 |
+
"num_tokens": 1214292.0,
|
| 950 |
+
"reward": 0.27426249384880064,
|
| 951 |
+
"reward_std": 0.42773920893669126,
|
| 952 |
+
"rewards/dbre_reward/mean": 0.27426249384880064,
|
| 953 |
+
"rewards/dbre_reward/std": 0.42773920893669126,
|
| 954 |
+
"step": 175,
|
| 955 |
+
"step_time": 32.60600872279319
|
| 956 |
+
},
|
| 957 |
+
{
|
| 958 |
+
"clip_ratio/high_max": 0.0,
|
| 959 |
+
"clip_ratio/high_mean": 0.0,
|
| 960 |
+
"clip_ratio/low_mean": 0.0,
|
| 961 |
+
"clip_ratio/low_min": 0.0,
|
| 962 |
+
"clip_ratio/region_mean": 0.0,
|
| 963 |
+
"completions/clipped_ratio": 0.9875,
|
| 964 |
+
"completions/max_length": 256.0,
|
| 965 |
+
"completions/max_terminated_length": 8.2,
|
| 966 |
+
"completions/mean_length": 253.3125,
|
| 967 |
+
"completions/mean_terminated_length": 8.2,
|
| 968 |
+
"completions/min_length": 213.0,
|
| 969 |
+
"completions/min_terminated_length": 8.2,
|
| 970 |
+
"entropy": 1.0290320612490178,
|
| 971 |
+
"epoch": 3.6,
|
| 972 |
+
"frac_reward_zero_std": 0.0,
|
| 973 |
+
"grad_norm": 0.09912109375,
|
| 974 |
+
"learning_rate": 3.21e-05,
|
| 975 |
+
"loss": -0.005978656560182571,
|
| 976 |
+
"num_tokens": 1249037.0,
|
| 977 |
+
"reward": 0.2384750008583069,
|
| 978 |
+
"reward_std": 0.4241094350814819,
|
| 979 |
+
"rewards/dbre_reward/mean": 0.2384750008583069,
|
| 980 |
+
"rewards/dbre_reward/std": 0.4241094350814819,
|
| 981 |
+
"step": 180,
|
| 982 |
+
"step_time": 33.03327611140848
|
| 983 |
+
},
|
| 984 |
+
{
|
| 985 |
+
"clip_ratio/high_max": 0.0,
|
| 986 |
+
"clip_ratio/high_mean": 0.0,
|
| 987 |
+
"clip_ratio/low_mean": 0.0,
|
| 988 |
+
"clip_ratio/low_min": 0.0,
|
| 989 |
+
"clip_ratio/region_mean": 0.0,
|
| 990 |
+
"completions/clipped_ratio": 0.925,
|
| 991 |
+
"completions/max_length": 256.0,
|
| 992 |
+
"completions/max_terminated_length": 132.8,
|
| 993 |
+
"completions/mean_length": 246.1875,
|
| 994 |
+
"completions/mean_terminated_length": 131.9,
|
| 995 |
+
"completions/min_length": 131.0,
|
| 996 |
+
"completions/min_terminated_length": 131.0,
|
| 997 |
+
"entropy": 0.92076805382967,
|
| 998 |
+
"epoch": 3.7,
|
| 999 |
+
"frac_reward_zero_std": 0.0,
|
| 1000 |
+
"grad_norm": 0.09619140625,
|
| 1001 |
+
"learning_rate": 3.16e-05,
|
| 1002 |
+
"loss": -0.009967343509197235,
|
| 1003 |
+
"num_tokens": 1283212.0,
|
| 1004 |
+
"reward": 0.33720000386238097,
|
| 1005 |
+
"reward_std": 0.46260204911231995,
|
| 1006 |
+
"rewards/dbre_reward/mean": 0.33720000386238097,
|
| 1007 |
+
"rewards/dbre_reward/std": 0.4626020550727844,
|
| 1008 |
+
"step": 185,
|
| 1009 |
+
"step_time": 34.82510126640555
|
| 1010 |
+
},
|
| 1011 |
+
{
|
| 1012 |
+
"clip_ratio/high_max": 0.0,
|
| 1013 |
+
"clip_ratio/high_mean": 0.0,
|
| 1014 |
+
"clip_ratio/low_mean": 0.0,
|
| 1015 |
+
"clip_ratio/low_min": 0.0,
|
| 1016 |
+
"clip_ratio/region_mean": 0.0,
|
| 1017 |
+
"completions/clipped_ratio": 0.95,
|
| 1018 |
+
"completions/max_length": 256.0,
|
| 1019 |
+
"completions/max_terminated_length": 74.2,
|
| 1020 |
+
"completions/mean_length": 248.875,
|
| 1021 |
+
"completions/mean_terminated_length": 45.4,
|
| 1022 |
+
"completions/min_length": 170.2,
|
| 1023 |
+
"completions/min_terminated_length": 16.6,
|
| 1024 |
+
"entropy": 1.0851470515131951,
|
| 1025 |
+
"epoch": 3.8,
|
| 1026 |
+
"frac_reward_zero_std": 0.2,
|
| 1027 |
+
"grad_norm": 0.09765625,
|
| 1028 |
+
"learning_rate": 3.1100000000000004e-05,
|
| 1029 |
+
"loss": 0.006586405634880066,
|
| 1030 |
+
"num_tokens": 1317602.0,
|
| 1031 |
+
"reward": 0.2771000027656555,
|
| 1032 |
+
"reward_std": 0.43828830122947693,
|
| 1033 |
+
"rewards/dbre_reward/mean": 0.2771000027656555,
|
| 1034 |
+
"rewards/dbre_reward/std": 0.43828831911087035,
|
| 1035 |
+
"step": 190,
|
| 1036 |
+
"step_time": 36.24654102979984
|
| 1037 |
+
},
|
| 1038 |
+
{
|
| 1039 |
+
"clip_ratio/high_max": 0.0,
|
| 1040 |
+
"clip_ratio/high_mean": 0.0,
|
| 1041 |
+
"clip_ratio/low_mean": 0.0,
|
| 1042 |
+
"clip_ratio/low_min": 0.0,
|
| 1043 |
+
"clip_ratio/region_mean": 0.0,
|
| 1044 |
+
"completions/clipped_ratio": 0.95,
|
| 1045 |
+
"completions/max_length": 256.0,
|
| 1046 |
+
"completions/max_terminated_length": 103.4,
|
| 1047 |
+
"completions/mean_length": 249.6625,
|
| 1048 |
+
"completions/mean_terminated_length": 103.4,
|
| 1049 |
+
"completions/min_length": 154.6,
|
| 1050 |
+
"completions/min_terminated_length": 103.4,
|
| 1051 |
+
"entropy": 0.9322984531521797,
|
| 1052 |
+
"epoch": 3.9,
|
| 1053 |
+
"frac_reward_zero_std": 0.0,
|
| 1054 |
+
"grad_norm": 0.099609375,
|
| 1055 |
+
"learning_rate": 3.06e-05,
|
| 1056 |
+
"loss": -0.011462598294019698,
|
| 1057 |
+
"num_tokens": 1352055.0,
|
| 1058 |
+
"reward": 0.30278749763965607,
|
| 1059 |
+
"reward_std": 0.4452593445777893,
|
| 1060 |
+
"rewards/dbre_reward/mean": 0.30278749763965607,
|
| 1061 |
+
"rewards/dbre_reward/std": 0.4452593445777893,
|
| 1062 |
+
"step": 195,
|
| 1063 |
+
"step_time": 29.027419628202914
|
| 1064 |
+
},
|
| 1065 |
+
{
|
| 1066 |
+
"clip_ratio/high_max": 0.0,
|
| 1067 |
+
"clip_ratio/high_mean": 0.0,
|
| 1068 |
+
"clip_ratio/low_mean": 0.0,
|
| 1069 |
+
"clip_ratio/low_min": 0.0,
|
| 1070 |
+
"clip_ratio/region_mean": 0.0,
|
| 1071 |
+
"completions/clipped_ratio": 0.975,
|
| 1072 |
+
"completions/max_length": 256.0,
|
| 1073 |
+
"completions/max_terminated_length": 21.8,
|
| 1074 |
+
"completions/mean_length": 250.9625,
|
| 1075 |
+
"completions/mean_terminated_length": 21.8,
|
| 1076 |
+
"completions/min_length": 175.4,
|
| 1077 |
+
"completions/min_terminated_length": 21.8,
|
| 1078 |
+
"entropy": 1.0117496035993099,
|
| 1079 |
+
"epoch": 4.0,
|
| 1080 |
+
"frac_reward_zero_std": 0.0,
|
| 1081 |
+
"grad_norm": 0.10009765625,
|
| 1082 |
+
"learning_rate": 3.01e-05,
|
| 1083 |
+
"loss": -0.017712239921092988,
|
| 1084 |
+
"num_tokens": 1386612.0,
|
| 1085 |
+
"reward": 0.2791749984025955,
|
| 1086 |
+
"reward_std": 0.43711588978767396,
|
| 1087 |
+
"rewards/dbre_reward/mean": 0.2791749984025955,
|
| 1088 |
+
"rewards/dbre_reward/std": 0.4371159017086029,
|
| 1089 |
+
"step": 200,
|
| 1090 |
+
"step_time": 27.529542883395333
|
| 1091 |
+
}
|
| 1092 |
+
],
|
| 1093 |
+
"logging_steps": 5,
|
| 1094 |
+
"max_steps": 500,
|
| 1095 |
+
"num_input_tokens_seen": 1386612,
|
| 1096 |
+
"num_train_epochs": 10,
|
| 1097 |
+
"save_steps": 50,
|
| 1098 |
+
"stateful_callbacks": {
|
| 1099 |
+
"TrainerControl": {
|
| 1100 |
+
"args": {
|
| 1101 |
+
"should_epoch_stop": false,
|
| 1102 |
+
"should_evaluate": false,
|
| 1103 |
+
"should_log": false,
|
| 1104 |
+
"should_save": true,
|
| 1105 |
+
"should_training_stop": false
|
| 1106 |
+
},
|
| 1107 |
+
"attributes": {}
|
| 1108 |
+
}
|
| 1109 |
+
},
|
| 1110 |
+
"total_flos": 0.0,
|
| 1111 |
+
"train_batch_size": 2,
|
| 1112 |
+
"trial_name": null,
|
| 1113 |
+
"trial_params": null
|
| 1114 |
+
}
|
grpo_dbre/checkpoint-200/training_args.bin
ADDED
|
Binary file (7.12 kB). View file
|
|
|
grpo_dbre/checkpoint-250/README.md
ADDED
|
@@ -0,0 +1,209 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
base_model: Qwen/Qwen2.5-Coder-1.5B-Instruct
|
| 3 |
+
library_name: peft
|
| 4 |
+
pipeline_tag: text-generation
|
| 5 |
+
tags:
|
| 6 |
+
- base_model:adapter:Qwen/Qwen2.5-Coder-1.5B-Instruct
|
| 7 |
+
- grpo
|
| 8 |
+
- lora
|
| 9 |
+
- transformers
|
| 10 |
+
- trl
|
| 11 |
+
---
|
| 12 |
+
|
| 13 |
+
# Model Card for Model ID
|
| 14 |
+
|
| 15 |
+
<!-- Provide a quick summary of what the model is/does. -->
|
| 16 |
+
|
| 17 |
+
|
| 18 |
+
|
| 19 |
+
## Model Details
|
| 20 |
+
|
| 21 |
+
### Model Description
|
| 22 |
+
|
| 23 |
+
<!-- Provide a longer summary of what this model is. -->
|
| 24 |
+
|
| 25 |
+
|
| 26 |
+
|
| 27 |
+
- **Developed by:** [More Information Needed]
|
| 28 |
+
- **Funded by [optional]:** [More Information Needed]
|
| 29 |
+
- **Shared by [optional]:** [More Information Needed]
|
| 30 |
+
- **Model type:** [More Information Needed]
|
| 31 |
+
- **Language(s) (NLP):** [More Information Needed]
|
| 32 |
+
- **License:** [More Information Needed]
|
| 33 |
+
- **Finetuned from model [optional]:** [More Information Needed]
|
| 34 |
+
|
| 35 |
+
### Model Sources [optional]
|
| 36 |
+
|
| 37 |
+
<!-- Provide the basic links for the model. -->
|
| 38 |
+
|
| 39 |
+
- **Repository:** [More Information Needed]
|
| 40 |
+
- **Paper [optional]:** [More Information Needed]
|
| 41 |
+
- **Demo [optional]:** [More Information Needed]
|
| 42 |
+
|
| 43 |
+
## Uses
|
| 44 |
+
|
| 45 |
+
<!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
|
| 46 |
+
|
| 47 |
+
### Direct Use
|
| 48 |
+
|
| 49 |
+
<!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
|
| 50 |
+
|
| 51 |
+
[More Information Needed]
|
| 52 |
+
|
| 53 |
+
### Downstream Use [optional]
|
| 54 |
+
|
| 55 |
+
<!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
|
| 56 |
+
|
| 57 |
+
[More Information Needed]
|
| 58 |
+
|
| 59 |
+
### Out-of-Scope Use
|
| 60 |
+
|
| 61 |
+
<!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
|
| 62 |
+
|
| 63 |
+
[More Information Needed]
|
| 64 |
+
|
| 65 |
+
## Bias, Risks, and Limitations
|
| 66 |
+
|
| 67 |
+
<!-- This section is meant to convey both technical and sociotechnical limitations. -->
|
| 68 |
+
|
| 69 |
+
[More Information Needed]
|
| 70 |
+
|
| 71 |
+
### Recommendations
|
| 72 |
+
|
| 73 |
+
<!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
|
| 74 |
+
|
| 75 |
+
Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
|
| 76 |
+
|
| 77 |
+
## How to Get Started with the Model
|
| 78 |
+
|
| 79 |
+
Use the code below to get started with the model.
|
| 80 |
+
|
| 81 |
+
[More Information Needed]
|
| 82 |
+
|
| 83 |
+
## Training Details
|
| 84 |
+
|
| 85 |
+
### Training Data
|
| 86 |
+
|
| 87 |
+
<!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
|
| 88 |
+
|
| 89 |
+
[More Information Needed]
|
| 90 |
+
|
| 91 |
+
### Training Procedure
|
| 92 |
+
|
| 93 |
+
<!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
|
| 94 |
+
|
| 95 |
+
#### Preprocessing [optional]
|
| 96 |
+
|
| 97 |
+
[More Information Needed]
|
| 98 |
+
|
| 99 |
+
|
| 100 |
+
#### Training Hyperparameters
|
| 101 |
+
|
| 102 |
+
- **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
|
| 103 |
+
|
| 104 |
+
#### Speeds, Sizes, Times [optional]
|
| 105 |
+
|
| 106 |
+
<!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
|
| 107 |
+
|
| 108 |
+
[More Information Needed]
|
| 109 |
+
|
| 110 |
+
## Evaluation
|
| 111 |
+
|
| 112 |
+
<!-- This section describes the evaluation protocols and provides the results. -->
|
| 113 |
+
|
| 114 |
+
### Testing Data, Factors & Metrics
|
| 115 |
+
|
| 116 |
+
#### Testing Data
|
| 117 |
+
|
| 118 |
+
<!-- This should link to a Dataset Card if possible. -->
|
| 119 |
+
|
| 120 |
+
[More Information Needed]
|
| 121 |
+
|
| 122 |
+
#### Factors
|
| 123 |
+
|
| 124 |
+
<!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
|
| 125 |
+
|
| 126 |
+
[More Information Needed]
|
| 127 |
+
|
| 128 |
+
#### Metrics
|
| 129 |
+
|
| 130 |
+
<!-- These are the evaluation metrics being used, ideally with a description of why. -->
|
| 131 |
+
|
| 132 |
+
[More Information Needed]
|
| 133 |
+
|
| 134 |
+
### Results
|
| 135 |
+
|
| 136 |
+
[More Information Needed]
|
| 137 |
+
|
| 138 |
+
#### Summary
|
| 139 |
+
|
| 140 |
+
|
| 141 |
+
|
| 142 |
+
## Model Examination [optional]
|
| 143 |
+
|
| 144 |
+
<!-- Relevant interpretability work for the model goes here -->
|
| 145 |
+
|
| 146 |
+
[More Information Needed]
|
| 147 |
+
|
| 148 |
+
## Environmental Impact
|
| 149 |
+
|
| 150 |
+
<!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
|
| 151 |
+
|
| 152 |
+
Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
|
| 153 |
+
|
| 154 |
+
- **Hardware Type:** [More Information Needed]
|
| 155 |
+
- **Hours used:** [More Information Needed]
|
| 156 |
+
- **Cloud Provider:** [More Information Needed]
|
| 157 |
+
- **Compute Region:** [More Information Needed]
|
| 158 |
+
- **Carbon Emitted:** [More Information Needed]
|
| 159 |
+
|
| 160 |
+
## Technical Specifications [optional]
|
| 161 |
+
|
| 162 |
+
### Model Architecture and Objective
|
| 163 |
+
|
| 164 |
+
[More Information Needed]
|
| 165 |
+
|
| 166 |
+
### Compute Infrastructure
|
| 167 |
+
|
| 168 |
+
[More Information Needed]
|
| 169 |
+
|
| 170 |
+
#### Hardware
|
| 171 |
+
|
| 172 |
+
[More Information Needed]
|
| 173 |
+
|
| 174 |
+
#### Software
|
| 175 |
+
|
| 176 |
+
[More Information Needed]
|
| 177 |
+
|
| 178 |
+
## Citation [optional]
|
| 179 |
+
|
| 180 |
+
<!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
|
| 181 |
+
|
| 182 |
+
**BibTeX:**
|
| 183 |
+
|
| 184 |
+
[More Information Needed]
|
| 185 |
+
|
| 186 |
+
**APA:**
|
| 187 |
+
|
| 188 |
+
[More Information Needed]
|
| 189 |
+
|
| 190 |
+
## Glossary [optional]
|
| 191 |
+
|
| 192 |
+
<!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
|
| 193 |
+
|
| 194 |
+
[More Information Needed]
|
| 195 |
+
|
| 196 |
+
## More Information [optional]
|
| 197 |
+
|
| 198 |
+
[More Information Needed]
|
| 199 |
+
|
| 200 |
+
## Model Card Authors [optional]
|
| 201 |
+
|
| 202 |
+
[More Information Needed]
|
| 203 |
+
|
| 204 |
+
## Model Card Contact
|
| 205 |
+
|
| 206 |
+
[More Information Needed]
|
| 207 |
+
### Framework versions
|
| 208 |
+
|
| 209 |
+
- PEFT 0.19.1
|
grpo_dbre/checkpoint-250/adapter_config.json
ADDED
|
@@ -0,0 +1,45 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"alora_invocation_tokens": null,
|
| 3 |
+
"alpha_pattern": {},
|
| 4 |
+
"arrow_config": null,
|
| 5 |
+
"auto_mapping": null,
|
| 6 |
+
"base_model_name_or_path": "Qwen/Qwen2.5-Coder-1.5B-Instruct",
|
| 7 |
+
"bias": "none",
|
| 8 |
+
"corda_config": null,
|
| 9 |
+
"ensure_weight_tying": false,
|
| 10 |
+
"eva_config": null,
|
| 11 |
+
"exclude_modules": null,
|
| 12 |
+
"fan_in_fan_out": false,
|
| 13 |
+
"inference_mode": true,
|
| 14 |
+
"init_lora_weights": true,
|
| 15 |
+
"layer_replication": null,
|
| 16 |
+
"layers_pattern": null,
|
| 17 |
+
"layers_to_transform": null,
|
| 18 |
+
"loftq_config": {},
|
| 19 |
+
"lora_alpha": 16,
|
| 20 |
+
"lora_bias": false,
|
| 21 |
+
"lora_dropout": 0,
|
| 22 |
+
"lora_ga_config": null,
|
| 23 |
+
"megatron_config": null,
|
| 24 |
+
"megatron_core": "megatron.core",
|
| 25 |
+
"modules_to_save": null,
|
| 26 |
+
"peft_type": "LORA",
|
| 27 |
+
"peft_version": "0.19.1",
|
| 28 |
+
"qalora_group_size": 16,
|
| 29 |
+
"r": 16,
|
| 30 |
+
"rank_pattern": {},
|
| 31 |
+
"revision": null,
|
| 32 |
+
"target_modules": [
|
| 33 |
+
"q_proj",
|
| 34 |
+
"k_proj",
|
| 35 |
+
"o_proj",
|
| 36 |
+
"v_proj"
|
| 37 |
+
],
|
| 38 |
+
"target_parameters": null,
|
| 39 |
+
"task_type": "CAUSAL_LM",
|
| 40 |
+
"trainable_token_indices": null,
|
| 41 |
+
"use_bdlora": null,
|
| 42 |
+
"use_dora": false,
|
| 43 |
+
"use_qalora": false,
|
| 44 |
+
"use_rslora": false
|
| 45 |
+
}
|
grpo_dbre/checkpoint-250/chat_template.jinja
ADDED
|
@@ -0,0 +1,54 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{%- if tools %}
|
| 2 |
+
{{- '<|im_start|>system\n' }}
|
| 3 |
+
{%- if messages[0]['role'] == 'system' %}
|
| 4 |
+
{{- messages[0]['content'] }}
|
| 5 |
+
{%- else %}
|
| 6 |
+
{{- 'You are Qwen, created by Alibaba Cloud. You are a helpful assistant.' }}
|
| 7 |
+
{%- endif %}
|
| 8 |
+
{{- "\n\n# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>" }}
|
| 9 |
+
{%- for tool in tools %}
|
| 10 |
+
{{- "\n" }}
|
| 11 |
+
{{- tool | tojson }}
|
| 12 |
+
{%- endfor %}
|
| 13 |
+
{{- "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call><|im_end|>\n" }}
|
| 14 |
+
{%- else %}
|
| 15 |
+
{%- if messages[0]['role'] == 'system' %}
|
| 16 |
+
{{- '<|im_start|>system\n' + messages[0]['content'] + '<|im_end|>\n' }}
|
| 17 |
+
{%- else %}
|
| 18 |
+
{{- '<|im_start|>system\nYou are Qwen, created by Alibaba Cloud. You are a helpful assistant.<|im_end|>\n' }}
|
| 19 |
+
{%- endif %}
|
| 20 |
+
{%- endif %}
|
| 21 |
+
{%- for message in messages %}
|
| 22 |
+
{%- if (message.role == "user") or (message.role == "system" and not loop.first) or (message.role == "assistant" and not message.tool_calls) %}
|
| 23 |
+
{{- '<|im_start|>' + message.role + '\n' + message.content + '<|im_end|>' + '\n' }}
|
| 24 |
+
{%- elif message.role == "assistant" %}
|
| 25 |
+
{{- '<|im_start|>' + message.role }}
|
| 26 |
+
{%- if message.content %}
|
| 27 |
+
{{- '\n' + message.content }}
|
| 28 |
+
{%- endif %}
|
| 29 |
+
{%- for tool_call in message.tool_calls %}
|
| 30 |
+
{%- if tool_call.function is defined %}
|
| 31 |
+
{%- set tool_call = tool_call.function %}
|
| 32 |
+
{%- endif %}
|
| 33 |
+
{{- '\n<tool_call>\n{"name": "' }}
|
| 34 |
+
{{- tool_call.name }}
|
| 35 |
+
{{- '", "arguments": ' }}
|
| 36 |
+
{{- tool_call.arguments | tojson }}
|
| 37 |
+
{{- '}\n</tool_call>' }}
|
| 38 |
+
{%- endfor %}
|
| 39 |
+
{{- '<|im_end|>\n' }}
|
| 40 |
+
{%- elif message.role == "tool" %}
|
| 41 |
+
{%- if (loop.index0 == 0) or (messages[loop.index0 - 1].role != "tool") %}
|
| 42 |
+
{{- '<|im_start|>user' }}
|
| 43 |
+
{%- endif %}
|
| 44 |
+
{{- '\n<tool_response>\n' }}
|
| 45 |
+
{{- message.content }}
|
| 46 |
+
{{- '\n</tool_response>' }}
|
| 47 |
+
{%- if loop.last or (messages[loop.index0 + 1].role != "tool") %}
|
| 48 |
+
{{- '<|im_end|>\n' }}
|
| 49 |
+
{%- endif %}
|
| 50 |
+
{%- endif %}
|
| 51 |
+
{%- endfor %}
|
| 52 |
+
{%- if add_generation_prompt %}
|
| 53 |
+
{{- '<|im_start|>assistant\n' }}
|
| 54 |
+
{%- endif %}
|
grpo_dbre/checkpoint-250/rng_state.pth
ADDED
|
Binary file (14.6 kB). View file
|
|
|
grpo_dbre/checkpoint-250/scheduler.pt
ADDED
|
Binary file (1.47 kB). View file
|
|
|
grpo_dbre/checkpoint-250/tokenizer_config.json
ADDED
|
@@ -0,0 +1,30 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"add_prefix_space": false,
|
| 3 |
+
"backend": "tokenizers",
|
| 4 |
+
"bos_token": null,
|
| 5 |
+
"clean_up_tokenization_spaces": false,
|
| 6 |
+
"eos_token": "<|im_end|>",
|
| 7 |
+
"errors": "replace",
|
| 8 |
+
"extra_special_tokens": [
|
| 9 |
+
"<|im_start|>",
|
| 10 |
+
"<|im_end|>",
|
| 11 |
+
"<|object_ref_start|>",
|
| 12 |
+
"<|object_ref_end|>",
|
| 13 |
+
"<|box_start|>",
|
| 14 |
+
"<|box_end|>",
|
| 15 |
+
"<|quad_start|>",
|
| 16 |
+
"<|quad_end|>",
|
| 17 |
+
"<|vision_start|>",
|
| 18 |
+
"<|vision_end|>",
|
| 19 |
+
"<|vision_pad|>",
|
| 20 |
+
"<|image_pad|>",
|
| 21 |
+
"<|video_pad|>"
|
| 22 |
+
],
|
| 23 |
+
"is_local": false,
|
| 24 |
+
"local_files_only": false,
|
| 25 |
+
"model_max_length": 32768,
|
| 26 |
+
"pad_token": "<|im_end|>",
|
| 27 |
+
"split_special_tokens": false,
|
| 28 |
+
"tokenizer_class": "Qwen2Tokenizer",
|
| 29 |
+
"unk_token": null
|
| 30 |
+
}
|
grpo_dbre/checkpoint-250/trainer_state.json
ADDED
|
@@ -0,0 +1,1384 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"best_global_step": null,
|
| 3 |
+
"best_metric": null,
|
| 4 |
+
"best_model_checkpoint": null,
|
| 5 |
+
"epoch": 5.0,
|
| 6 |
+
"eval_steps": 500,
|
| 7 |
+
"global_step": 250,
|
| 8 |
+
"is_hyper_param_search": false,
|
| 9 |
+
"is_local_process_zero": true,
|
| 10 |
+
"is_world_process_zero": true,
|
| 11 |
+
"log_history": [
|
| 12 |
+
{
|
| 13 |
+
"clip_ratio/high_max": 0.0,
|
| 14 |
+
"clip_ratio/high_mean": 0.0,
|
| 15 |
+
"clip_ratio/low_mean": 0.0,
|
| 16 |
+
"clip_ratio/low_min": 0.0,
|
| 17 |
+
"clip_ratio/region_mean": 0.0,
|
| 18 |
+
"completions/clipped_ratio": 0.9875,
|
| 19 |
+
"completions/max_length": 256.0,
|
| 20 |
+
"completions/max_terminated_length": 19.6,
|
| 21 |
+
"completions/mean_length": 254.025,
|
| 22 |
+
"completions/mean_terminated_length": 19.6,
|
| 23 |
+
"completions/min_length": 224.4,
|
| 24 |
+
"completions/min_terminated_length": 19.6,
|
| 25 |
+
"entropy": 0.903753462433815,
|
| 26 |
+
"epoch": 0.1,
|
| 27 |
+
"frac_reward_zero_std": 0.8,
|
| 28 |
+
"grad_norm": 0.046875,
|
| 29 |
+
"learning_rate": 4.96e-05,
|
| 30 |
+
"loss": -2.2351741790771484e-09,
|
| 31 |
+
"num_tokens": 34802.0,
|
| 32 |
+
"reward": 0.025,
|
| 33 |
+
"reward_std": 0.1,
|
| 34 |
+
"rewards/dbre_reward/mean": 0.025,
|
| 35 |
+
"rewards/dbre_reward/std": 0.1,
|
| 36 |
+
"step": 5,
|
| 37 |
+
"step_time": 27.396174477002933
|
| 38 |
+
},
|
| 39 |
+
{
|
| 40 |
+
"clip_ratio/high_max": 0.0,
|
| 41 |
+
"clip_ratio/high_mean": 0.0,
|
| 42 |
+
"clip_ratio/low_mean": 0.0,
|
| 43 |
+
"clip_ratio/low_min": 0.0,
|
| 44 |
+
"clip_ratio/region_mean": 0.0,
|
| 45 |
+
"completions/clipped_ratio": 0.975,
|
| 46 |
+
"completions/max_length": 256.0,
|
| 47 |
+
"completions/max_terminated_length": 48.4,
|
| 48 |
+
"completions/mean_length": 252.625,
|
| 49 |
+
"completions/mean_terminated_length": 48.4,
|
| 50 |
+
"completions/min_length": 202.0,
|
| 51 |
+
"completions/min_terminated_length": 48.4,
|
| 52 |
+
"entropy": 0.9422640666365624,
|
| 53 |
+
"epoch": 0.2,
|
| 54 |
+
"frac_reward_zero_std": 0.5,
|
| 55 |
+
"grad_norm": 0.0673828125,
|
| 56 |
+
"learning_rate": 4.91e-05,
|
| 57 |
+
"loss": -5.960464477539063e-09,
|
| 58 |
+
"num_tokens": 69492.0,
|
| 59 |
+
"reward": 0.09868749976158142,
|
| 60 |
+
"reward_std": 0.25848535895347596,
|
| 61 |
+
"rewards/dbre_reward/mean": 0.09868749976158142,
|
| 62 |
+
"rewards/dbre_reward/std": 0.2584853649139404,
|
| 63 |
+
"step": 10,
|
| 64 |
+
"step_time": 27.675077842207976
|
| 65 |
+
},
|
| 66 |
+
{
|
| 67 |
+
"clip_ratio/high_max": 0.0,
|
| 68 |
+
"clip_ratio/high_mean": 0.0,
|
| 69 |
+
"clip_ratio/low_mean": 0.0,
|
| 70 |
+
"clip_ratio/low_min": 0.0,
|
| 71 |
+
"clip_ratio/region_mean": 0.0,
|
| 72 |
+
"completions/clipped_ratio": 0.975,
|
| 73 |
+
"completions/max_length": 256.0,
|
| 74 |
+
"completions/max_terminated_length": 79.0,
|
| 75 |
+
"completions/mean_length": 254.5375,
|
| 76 |
+
"completions/mean_terminated_length": 79.0,
|
| 77 |
+
"completions/min_length": 232.6,
|
| 78 |
+
"completions/min_terminated_length": 79.0,
|
| 79 |
+
"entropy": 0.9289400212466716,
|
| 80 |
+
"epoch": 0.3,
|
| 81 |
+
"frac_reward_zero_std": 0.5,
|
| 82 |
+
"grad_norm": 0.059326171875,
|
| 83 |
+
"learning_rate": 4.86e-05,
|
| 84 |
+
"loss": -0.0020709306001663206,
|
| 85 |
+
"num_tokens": 104335.0,
|
| 86 |
+
"reward": 0.11038749814033508,
|
| 87 |
+
"reward_std": 0.27505697011947633,
|
| 88 |
+
"rewards/dbre_reward/mean": 0.11038749814033508,
|
| 89 |
+
"rewards/dbre_reward/std": 0.2750569820404053,
|
| 90 |
+
"step": 15,
|
| 91 |
+
"step_time": 27.678717108402633
|
| 92 |
+
},
|
| 93 |
+
{
|
| 94 |
+
"clip_ratio/high_max": 0.0,
|
| 95 |
+
"clip_ratio/high_mean": 0.0,
|
| 96 |
+
"clip_ratio/low_mean": 0.0,
|
| 97 |
+
"clip_ratio/low_min": 0.0,
|
| 98 |
+
"clip_ratio/region_mean": 0.0,
|
| 99 |
+
"completions/clipped_ratio": 0.9625,
|
| 100 |
+
"completions/max_length": 256.0,
|
| 101 |
+
"completions/max_terminated_length": 61.4,
|
| 102 |
+
"completions/mean_length": 250.2375,
|
| 103 |
+
"completions/mean_terminated_length": 61.4,
|
| 104 |
+
"completions/min_length": 163.8,
|
| 105 |
+
"completions/min_terminated_length": 61.4,
|
| 106 |
+
"entropy": 0.8890757068991662,
|
| 107 |
+
"epoch": 0.4,
|
| 108 |
+
"frac_reward_zero_std": 0.3,
|
| 109 |
+
"grad_norm": 0.052001953125,
|
| 110 |
+
"learning_rate": 4.8100000000000004e-05,
|
| 111 |
+
"loss": -0.00669153705239296,
|
| 112 |
+
"num_tokens": 138834.0,
|
| 113 |
+
"reward": 0.11190000027418137,
|
| 114 |
+
"reward_std": 0.3216355323791504,
|
| 115 |
+
"rewards/dbre_reward/mean": 0.11190000027418137,
|
| 116 |
+
"rewards/dbre_reward/std": 0.3216355502605438,
|
| 117 |
+
"step": 20,
|
| 118 |
+
"step_time": 27.725935825207852
|
| 119 |
+
},
|
| 120 |
+
{
|
| 121 |
+
"clip_ratio/high_max": 0.0,
|
| 122 |
+
"clip_ratio/high_mean": 0.0,
|
| 123 |
+
"clip_ratio/low_mean": 0.0,
|
| 124 |
+
"clip_ratio/low_min": 0.0,
|
| 125 |
+
"clip_ratio/region_mean": 0.0,
|
| 126 |
+
"completions/clipped_ratio": 0.975,
|
| 127 |
+
"completions/max_length": 256.0,
|
| 128 |
+
"completions/max_terminated_length": 78.2,
|
| 129 |
+
"completions/mean_length": 254.4875,
|
| 130 |
+
"completions/mean_terminated_length": 78.2,
|
| 131 |
+
"completions/min_length": 231.8,
|
| 132 |
+
"completions/min_terminated_length": 78.2,
|
| 133 |
+
"entropy": 0.906505486369133,
|
| 134 |
+
"epoch": 0.5,
|
| 135 |
+
"frac_reward_zero_std": 0.7,
|
| 136 |
+
"grad_norm": 0.0,
|
| 137 |
+
"learning_rate": 4.76e-05,
|
| 138 |
+
"loss": 0.00735630989074707,
|
| 139 |
+
"num_tokens": 173673.0,
|
| 140 |
+
"reward": 0.0375,
|
| 141 |
+
"reward_std": 0.15,
|
| 142 |
+
"rewards/dbre_reward/mean": 0.0375,
|
| 143 |
+
"rewards/dbre_reward/std": 0.15,
|
| 144 |
+
"step": 25,
|
| 145 |
+
"step_time": 27.84102148480888
|
| 146 |
+
},
|
| 147 |
+
{
|
| 148 |
+
"clip_ratio/high_max": 0.0,
|
| 149 |
+
"clip_ratio/high_mean": 0.0,
|
| 150 |
+
"clip_ratio/low_mean": 0.0,
|
| 151 |
+
"clip_ratio/low_min": 0.0,
|
| 152 |
+
"clip_ratio/region_mean": 0.0,
|
| 153 |
+
"completions/clipped_ratio": 0.9625,
|
| 154 |
+
"completions/max_length": 256.0,
|
| 155 |
+
"completions/max_terminated_length": 69.0,
|
| 156 |
+
"completions/mean_length": 251.2125,
|
| 157 |
+
"completions/mean_terminated_length": 56.1,
|
| 158 |
+
"completions/min_length": 196.8,
|
| 159 |
+
"completions/min_terminated_length": 43.2,
|
| 160 |
+
"entropy": 0.8586828224360943,
|
| 161 |
+
"epoch": 0.6,
|
| 162 |
+
"frac_reward_zero_std": 0.6,
|
| 163 |
+
"grad_norm": 0.059814453125,
|
| 164 |
+
"learning_rate": 4.71e-05,
|
| 165 |
+
"loss": -0.005647056177258492,
|
| 166 |
+
"num_tokens": 208250.0,
|
| 167 |
+
"reward": 0.0625,
|
| 168 |
+
"reward_std": 0.18662600517272948,
|
| 169 |
+
"rewards/dbre_reward/mean": 0.0625,
|
| 170 |
+
"rewards/dbre_reward/std": 0.18662601709365845,
|
| 171 |
+
"step": 30,
|
| 172 |
+
"step_time": 37.554382849001556
|
| 173 |
+
},
|
| 174 |
+
{
|
| 175 |
+
"clip_ratio/high_max": 0.0,
|
| 176 |
+
"clip_ratio/high_mean": 0.0,
|
| 177 |
+
"clip_ratio/low_mean": 0.0,
|
| 178 |
+
"clip_ratio/low_min": 0.0,
|
| 179 |
+
"clip_ratio/region_mean": 0.0,
|
| 180 |
+
"completions/clipped_ratio": 0.9625,
|
| 181 |
+
"completions/max_length": 256.0,
|
| 182 |
+
"completions/max_terminated_length": 115.6,
|
| 183 |
+
"completions/mean_length": 253.625,
|
| 184 |
+
"completions/mean_terminated_length": 115.6,
|
| 185 |
+
"completions/min_length": 218.0,
|
| 186 |
+
"completions/min_terminated_length": 115.6,
|
| 187 |
+
"entropy": 0.8725819021463395,
|
| 188 |
+
"epoch": 0.7,
|
| 189 |
+
"frac_reward_zero_std": 0.3,
|
| 190 |
+
"grad_norm": 0.06787109375,
|
| 191 |
+
"learning_rate": 4.660000000000001e-05,
|
| 192 |
+
"loss": 0.0011730872094631196,
|
| 193 |
+
"num_tokens": 243020.0,
|
| 194 |
+
"reward": 0.11033750027418136,
|
| 195 |
+
"reward_std": 0.31098498702049254,
|
| 196 |
+
"rewards/dbre_reward/mean": 0.11033750027418136,
|
| 197 |
+
"rewards/dbre_reward/std": 0.31098498702049254,
|
| 198 |
+
"step": 35,
|
| 199 |
+
"step_time": 35.878880111602484
|
| 200 |
+
},
|
| 201 |
+
{
|
| 202 |
+
"clip_ratio/high_max": 0.0,
|
| 203 |
+
"clip_ratio/high_mean": 0.0,
|
| 204 |
+
"clip_ratio/low_mean": 0.0,
|
| 205 |
+
"clip_ratio/low_min": 0.0,
|
| 206 |
+
"clip_ratio/region_mean": 0.0,
|
| 207 |
+
"completions/clipped_ratio": 1.0,
|
| 208 |
+
"completions/max_length": 256.0,
|
| 209 |
+
"completions/max_terminated_length": 0.0,
|
| 210 |
+
"completions/mean_length": 256.0,
|
| 211 |
+
"completions/mean_terminated_length": 0.0,
|
| 212 |
+
"completions/min_length": 256.0,
|
| 213 |
+
"completions/min_terminated_length": 0.0,
|
| 214 |
+
"entropy": 0.9090480573475361,
|
| 215 |
+
"epoch": 0.8,
|
| 216 |
+
"frac_reward_zero_std": 0.4,
|
| 217 |
+
"grad_norm": 0.046875,
|
| 218 |
+
"learning_rate": 4.61e-05,
|
| 219 |
+
"loss": -8.940696716308593e-09,
|
| 220 |
+
"num_tokens": 277980.0,
|
| 221 |
+
"reward": 0.09995000064373016,
|
| 222 |
+
"reward_std": 0.30480254292488096,
|
| 223 |
+
"rewards/dbre_reward/mean": 0.09995000064373016,
|
| 224 |
+
"rewards/dbre_reward/std": 0.3048025548458099,
|
| 225 |
+
"step": 40,
|
| 226 |
+
"step_time": 34.05516860120406
|
| 227 |
+
},
|
| 228 |
+
{
|
| 229 |
+
"clip_ratio/high_max": 0.0,
|
| 230 |
+
"clip_ratio/high_mean": 0.0,
|
| 231 |
+
"clip_ratio/low_mean": 0.0,
|
| 232 |
+
"clip_ratio/low_min": 0.0,
|
| 233 |
+
"clip_ratio/region_mean": 0.0,
|
| 234 |
+
"completions/clipped_ratio": 0.9875,
|
| 235 |
+
"completions/max_length": 256.0,
|
| 236 |
+
"completions/max_terminated_length": 26.0,
|
| 237 |
+
"completions/mean_length": 254.425,
|
| 238 |
+
"completions/mean_terminated_length": 26.0,
|
| 239 |
+
"completions/min_length": 230.8,
|
| 240 |
+
"completions/min_terminated_length": 26.0,
|
| 241 |
+
"entropy": 0.8977848328649998,
|
| 242 |
+
"epoch": 0.9,
|
| 243 |
+
"frac_reward_zero_std": 0.6,
|
| 244 |
+
"grad_norm": 0.0,
|
| 245 |
+
"learning_rate": 4.5600000000000004e-05,
|
| 246 |
+
"loss": 1.4901161193847657e-09,
|
| 247 |
+
"num_tokens": 312814.0,
|
| 248 |
+
"reward": 0.07415000200271607,
|
| 249 |
+
"reward_std": 0.197099506855011,
|
| 250 |
+
"rewards/dbre_reward/mean": 0.07415000200271607,
|
| 251 |
+
"rewards/dbre_reward/std": 0.197099506855011,
|
| 252 |
+
"step": 45,
|
| 253 |
+
"step_time": 34.819741847200206
|
| 254 |
+
},
|
| 255 |
+
{
|
| 256 |
+
"clip_ratio/high_max": 0.0,
|
| 257 |
+
"clip_ratio/high_mean": 0.0,
|
| 258 |
+
"clip_ratio/low_mean": 0.0,
|
| 259 |
+
"clip_ratio/low_min": 0.0,
|
| 260 |
+
"clip_ratio/region_mean": 0.0,
|
| 261 |
+
"completions/clipped_ratio": 0.975,
|
| 262 |
+
"completions/max_length": 256.0,
|
| 263 |
+
"completions/max_terminated_length": 20.4,
|
| 264 |
+
"completions/mean_length": 250.875,
|
| 265 |
+
"completions/mean_terminated_length": 20.4,
|
| 266 |
+
"completions/min_length": 174.0,
|
| 267 |
+
"completions/min_terminated_length": 20.4,
|
| 268 |
+
"entropy": 1.0103468239307403,
|
| 269 |
+
"epoch": 1.0,
|
| 270 |
+
"frac_reward_zero_std": 0.8,
|
| 271 |
+
"grad_norm": 0.0,
|
| 272 |
+
"learning_rate": 4.5100000000000005e-05,
|
| 273 |
+
"loss": -2.2351741790771484e-09,
|
| 274 |
+
"num_tokens": 347364.0,
|
| 275 |
+
"reward": 0.04707500040531158,
|
| 276 |
+
"reward_std": 0.12459058165550232,
|
| 277 |
+
"rewards/dbre_reward/mean": 0.04707500040531158,
|
| 278 |
+
"rewards/dbre_reward/std": 0.12459058165550232,
|
| 279 |
+
"step": 50,
|
| 280 |
+
"step_time": 36.14326306019211
|
| 281 |
+
},
|
| 282 |
+
{
|
| 283 |
+
"clip_ratio/high_max": 0.0,
|
| 284 |
+
"clip_ratio/high_mean": 0.0,
|
| 285 |
+
"clip_ratio/low_mean": 0.0,
|
| 286 |
+
"clip_ratio/low_min": 0.0,
|
| 287 |
+
"clip_ratio/region_mean": 0.0,
|
| 288 |
+
"completions/clipped_ratio": 0.9375,
|
| 289 |
+
"completions/max_length": 256.0,
|
| 290 |
+
"completions/max_terminated_length": 41.8,
|
| 291 |
+
"completions/mean_length": 244.5625,
|
| 292 |
+
"completions/mean_terminated_length": 18.55,
|
| 293 |
+
"completions/min_length": 161.4,
|
| 294 |
+
"completions/min_terminated_length": 7.8,
|
| 295 |
+
"entropy": 0.9930311724543571,
|
| 296 |
+
"epoch": 1.1,
|
| 297 |
+
"frac_reward_zero_std": 0.6,
|
| 298 |
+
"grad_norm": 0.06396484375,
|
| 299 |
+
"learning_rate": 4.46e-05,
|
| 300 |
+
"loss": -0.017940016090869905,
|
| 301 |
+
"num_tokens": 381409.0,
|
| 302 |
+
"reward": 0.06066250056028366,
|
| 303 |
+
"reward_std": 0.2124839812517166,
|
| 304 |
+
"rewards/dbre_reward/mean": 0.06066250056028366,
|
| 305 |
+
"rewards/dbre_reward/std": 0.21248398423194886,
|
| 306 |
+
"step": 55,
|
| 307 |
+
"step_time": 28.94829335878603
|
| 308 |
+
},
|
| 309 |
+
{
|
| 310 |
+
"clip_ratio/high_max": 0.0,
|
| 311 |
+
"clip_ratio/high_mean": 0.0,
|
| 312 |
+
"clip_ratio/low_mean": 0.0,
|
| 313 |
+
"clip_ratio/low_min": 0.0,
|
| 314 |
+
"clip_ratio/region_mean": 0.0,
|
| 315 |
+
"completions/clipped_ratio": 0.9875,
|
| 316 |
+
"completions/max_length": 256.0,
|
| 317 |
+
"completions/max_terminated_length": 9.6,
|
| 318 |
+
"completions/mean_length": 253.4,
|
| 319 |
+
"completions/mean_terminated_length": 9.6,
|
| 320 |
+
"completions/min_length": 214.4,
|
| 321 |
+
"completions/min_terminated_length": 9.6,
|
| 322 |
+
"entropy": 0.8150956228375434,
|
| 323 |
+
"epoch": 1.2,
|
| 324 |
+
"frac_reward_zero_std": 0.5,
|
| 325 |
+
"grad_norm": 0.06103515625,
|
| 326 |
+
"learning_rate": 4.41e-05,
|
| 327 |
+
"loss": -4.470348358154297e-09,
|
| 328 |
+
"num_tokens": 416161.0,
|
| 329 |
+
"reward": 0.11134999990463257,
|
| 330 |
+
"reward_std": 0.2740528523921967,
|
| 331 |
+
"rewards/dbre_reward/mean": 0.11134999990463257,
|
| 332 |
+
"rewards/dbre_reward/std": 0.27405285835266113,
|
| 333 |
+
"step": 60,
|
| 334 |
+
"step_time": 27.705245727201692
|
| 335 |
+
},
|
| 336 |
+
{
|
| 337 |
+
"clip_ratio/high_max": 0.0,
|
| 338 |
+
"clip_ratio/high_mean": 0.0,
|
| 339 |
+
"clip_ratio/low_mean": 0.0,
|
| 340 |
+
"clip_ratio/low_min": 0.0,
|
| 341 |
+
"clip_ratio/region_mean": 0.0,
|
| 342 |
+
"completions/clipped_ratio": 0.975,
|
| 343 |
+
"completions/max_length": 256.0,
|
| 344 |
+
"completions/max_terminated_length": 41.0,
|
| 345 |
+
"completions/mean_length": 252.1625,
|
| 346 |
+
"completions/mean_terminated_length": 41.0,
|
| 347 |
+
"completions/min_length": 194.6,
|
| 348 |
+
"completions/min_terminated_length": 41.0,
|
| 349 |
+
"entropy": 0.9309644259512424,
|
| 350 |
+
"epoch": 1.3,
|
| 351 |
+
"frac_reward_zero_std": 0.7,
|
| 352 |
+
"grad_norm": 0.0,
|
| 353 |
+
"learning_rate": 4.36e-05,
|
| 354 |
+
"loss": -8.940696716308593e-09,
|
| 355 |
+
"num_tokens": 450814.0,
|
| 356 |
+
"reward": 0.0875,
|
| 357 |
+
"reward_std": 0.21124515533447266,
|
| 358 |
+
"rewards/dbre_reward/mean": 0.0875,
|
| 359 |
+
"rewards/dbre_reward/std": 0.21124515533447266,
|
| 360 |
+
"step": 65,
|
| 361 |
+
"step_time": 27.76657635839365
|
| 362 |
+
},
|
| 363 |
+
{
|
| 364 |
+
"clip_ratio/high_max": 0.0,
|
| 365 |
+
"clip_ratio/high_mean": 0.0,
|
| 366 |
+
"clip_ratio/low_mean": 0.0,
|
| 367 |
+
"clip_ratio/low_min": 0.0,
|
| 368 |
+
"clip_ratio/region_mean": 0.0,
|
| 369 |
+
"completions/clipped_ratio": 0.9875,
|
| 370 |
+
"completions/max_length": 256.0,
|
| 371 |
+
"completions/max_terminated_length": 30.4,
|
| 372 |
+
"completions/mean_length": 254.7,
|
| 373 |
+
"completions/mean_terminated_length": 30.4,
|
| 374 |
+
"completions/min_length": 235.2,
|
| 375 |
+
"completions/min_terminated_length": 30.4,
|
| 376 |
+
"entropy": 0.8473479233682155,
|
| 377 |
+
"epoch": 1.4,
|
| 378 |
+
"frac_reward_zero_std": 0.6,
|
| 379 |
+
"grad_norm": 0.05078125,
|
| 380 |
+
"learning_rate": 4.3100000000000004e-05,
|
| 381 |
+
"loss": -2.2351741790771484e-09,
|
| 382 |
+
"num_tokens": 485670.0,
|
| 383 |
+
"reward": 0.07392499968409538,
|
| 384 |
+
"reward_std": 0.1946355789899826,
|
| 385 |
+
"rewards/dbre_reward/mean": 0.07392499968409538,
|
| 386 |
+
"rewards/dbre_reward/std": 0.1946355879306793,
|
| 387 |
+
"step": 70,
|
| 388 |
+
"step_time": 27.701584570185513
|
| 389 |
+
},
|
| 390 |
+
{
|
| 391 |
+
"clip_ratio/high_max": 0.0,
|
| 392 |
+
"clip_ratio/high_mean": 0.0,
|
| 393 |
+
"clip_ratio/low_mean": 0.0,
|
| 394 |
+
"clip_ratio/low_min": 0.0,
|
| 395 |
+
"clip_ratio/region_mean": 0.0,
|
| 396 |
+
"completions/clipped_ratio": 0.975,
|
| 397 |
+
"completions/max_length": 256.0,
|
| 398 |
+
"completions/max_terminated_length": 51.0,
|
| 399 |
+
"completions/mean_length": 252.7875,
|
| 400 |
+
"completions/mean_terminated_length": 51.0,
|
| 401 |
+
"completions/min_length": 204.6,
|
| 402 |
+
"completions/min_terminated_length": 51.0,
|
| 403 |
+
"entropy": 0.8822006396949291,
|
| 404 |
+
"epoch": 1.5,
|
| 405 |
+
"frac_reward_zero_std": 0.4,
|
| 406 |
+
"grad_norm": 0.046875,
|
| 407 |
+
"learning_rate": 4.26e-05,
|
| 408 |
+
"loss": -0.004587128758430481,
|
| 409 |
+
"num_tokens": 520373.0,
|
| 410 |
+
"reward": 0.09866249859333039,
|
| 411 |
+
"reward_std": 0.29479086995124815,
|
| 412 |
+
"rewards/dbre_reward/mean": 0.09866249859333039,
|
| 413 |
+
"rewards/dbre_reward/std": 0.29479087591171266,
|
| 414 |
+
"step": 75,
|
| 415 |
+
"step_time": 27.723117466596886
|
| 416 |
+
},
|
| 417 |
+
{
|
| 418 |
+
"clip_ratio/high_max": 0.0,
|
| 419 |
+
"clip_ratio/high_mean": 0.0,
|
| 420 |
+
"clip_ratio/low_mean": 0.0,
|
| 421 |
+
"clip_ratio/low_min": 0.0,
|
| 422 |
+
"clip_ratio/region_mean": 0.0,
|
| 423 |
+
"completions/clipped_ratio": 1.0,
|
| 424 |
+
"completions/max_length": 256.0,
|
| 425 |
+
"completions/max_terminated_length": 0.0,
|
| 426 |
+
"completions/mean_length": 256.0,
|
| 427 |
+
"completions/mean_terminated_length": 0.0,
|
| 428 |
+
"completions/min_length": 256.0,
|
| 429 |
+
"completions/min_terminated_length": 0.0,
|
| 430 |
+
"entropy": 0.8492010429501533,
|
| 431 |
+
"epoch": 1.6,
|
| 432 |
+
"frac_reward_zero_std": 0.4,
|
| 433 |
+
"grad_norm": 0.052734375,
|
| 434 |
+
"learning_rate": 4.21e-05,
|
| 435 |
+
"loss": -4.470348358154297e-09,
|
| 436 |
+
"num_tokens": 555333.0,
|
| 437 |
+
"reward": 0.13631249964237213,
|
| 438 |
+
"reward_std": 0.3453687012195587,
|
| 439 |
+
"rewards/dbre_reward/mean": 0.13631249964237213,
|
| 440 |
+
"rewards/dbre_reward/std": 0.34536872506141664,
|
| 441 |
+
"step": 80,
|
| 442 |
+
"step_time": 27.74355768300593
|
| 443 |
+
},
|
| 444 |
+
{
|
| 445 |
+
"clip_ratio/high_max": 0.0,
|
| 446 |
+
"clip_ratio/high_mean": 0.0,
|
| 447 |
+
"clip_ratio/low_mean": 0.0,
|
| 448 |
+
"clip_ratio/low_min": 0.0,
|
| 449 |
+
"clip_ratio/region_mean": 0.0,
|
| 450 |
+
"completions/clipped_ratio": 0.9875,
|
| 451 |
+
"completions/max_length": 256.0,
|
| 452 |
+
"completions/max_terminated_length": 23.0,
|
| 453 |
+
"completions/mean_length": 254.2375,
|
| 454 |
+
"completions/mean_terminated_length": 23.0,
|
| 455 |
+
"completions/min_length": 227.8,
|
| 456 |
+
"completions/min_terminated_length": 23.0,
|
| 457 |
+
"entropy": 0.9008926346898078,
|
| 458 |
+
"epoch": 1.7,
|
| 459 |
+
"frac_reward_zero_std": 0.4,
|
| 460 |
+
"grad_norm": 0.053466796875,
|
| 461 |
+
"learning_rate": 4.16e-05,
|
| 462 |
+
"loss": -0.0038484178483486177,
|
| 463 |
+
"num_tokens": 590152.0,
|
| 464 |
+
"reward": 0.13577499985694885,
|
| 465 |
+
"reward_std": 0.3495619535446167,
|
| 466 |
+
"rewards/dbre_reward/mean": 0.13577499985694885,
|
| 467 |
+
"rewards/dbre_reward/std": 0.34956197142601014,
|
| 468 |
+
"step": 85,
|
| 469 |
+
"step_time": 27.809819040997535
|
| 470 |
+
},
|
| 471 |
+
{
|
| 472 |
+
"clip_ratio/high_max": 0.0,
|
| 473 |
+
"clip_ratio/high_mean": 0.0,
|
| 474 |
+
"clip_ratio/low_mean": 0.0,
|
| 475 |
+
"clip_ratio/low_min": 0.0,
|
| 476 |
+
"clip_ratio/region_mean": 0.0,
|
| 477 |
+
"completions/clipped_ratio": 0.975,
|
| 478 |
+
"completions/max_length": 256.0,
|
| 479 |
+
"completions/max_terminated_length": 34.4,
|
| 480 |
+
"completions/mean_length": 253.7375,
|
| 481 |
+
"completions/mean_terminated_length": 33.1,
|
| 482 |
+
"completions/min_length": 236.6,
|
| 483 |
+
"completions/min_terminated_length": 31.8,
|
| 484 |
+
"entropy": 0.975049901008606,
|
| 485 |
+
"epoch": 1.8,
|
| 486 |
+
"frac_reward_zero_std": 0.3,
|
| 487 |
+
"grad_norm": 0.08447265625,
|
| 488 |
+
"learning_rate": 4.11e-05,
|
| 489 |
+
"loss": -0.0023170128464698792,
|
| 490 |
+
"num_tokens": 624931.0,
|
| 491 |
+
"reward": 0.14498749673366546,
|
| 492 |
+
"reward_std": 0.3355918139219284,
|
| 493 |
+
"rewards/dbre_reward/mean": 0.14498749673366546,
|
| 494 |
+
"rewards/dbre_reward/std": 0.3355918198823929,
|
| 495 |
+
"step": 90,
|
| 496 |
+
"step_time": 27.872812610585243
|
| 497 |
+
},
|
| 498 |
+
{
|
| 499 |
+
"clip_ratio/high_max": 0.0,
|
| 500 |
+
"clip_ratio/high_mean": 0.0,
|
| 501 |
+
"clip_ratio/low_mean": 0.0,
|
| 502 |
+
"clip_ratio/low_min": 0.0,
|
| 503 |
+
"clip_ratio/region_mean": 0.0,
|
| 504 |
+
"completions/clipped_ratio": 0.95,
|
| 505 |
+
"completions/max_length": 256.0,
|
| 506 |
+
"completions/max_terminated_length": 75.6,
|
| 507 |
+
"completions/mean_length": 250.775,
|
| 508 |
+
"completions/mean_terminated_length": 60.6,
|
| 509 |
+
"completions/min_length": 199.2,
|
| 510 |
+
"completions/min_terminated_length": 45.6,
|
| 511 |
+
"entropy": 0.9378853186964988,
|
| 512 |
+
"epoch": 1.9,
|
| 513 |
+
"frac_reward_zero_std": 0.4,
|
| 514 |
+
"grad_norm": 0.0576171875,
|
| 515 |
+
"learning_rate": 4.0600000000000004e-05,
|
| 516 |
+
"loss": -0.010608357191085816,
|
| 517 |
+
"num_tokens": 659473.0,
|
| 518 |
+
"reward": 0.13513749986886978,
|
| 519 |
+
"reward_std": 0.3332351267337799,
|
| 520 |
+
"rewards/dbre_reward/mean": 0.13513749986886978,
|
| 521 |
+
"rewards/dbre_reward/std": 0.3332351326942444,
|
| 522 |
+
"step": 95,
|
| 523 |
+
"step_time": 27.71230500699894
|
| 524 |
+
},
|
| 525 |
+
{
|
| 526 |
+
"clip_ratio/high_max": 0.0,
|
| 527 |
+
"clip_ratio/high_mean": 0.0,
|
| 528 |
+
"clip_ratio/low_mean": 0.0,
|
| 529 |
+
"clip_ratio/low_min": 0.0,
|
| 530 |
+
"clip_ratio/region_mean": 0.0,
|
| 531 |
+
"completions/clipped_ratio": 0.975,
|
| 532 |
+
"completions/max_length": 256.0,
|
| 533 |
+
"completions/max_terminated_length": 80.0,
|
| 534 |
+
"completions/mean_length": 254.6,
|
| 535 |
+
"completions/mean_terminated_length": 80.0,
|
| 536 |
+
"completions/min_length": 233.6,
|
| 537 |
+
"completions/min_terminated_length": 80.0,
|
| 538 |
+
"entropy": 0.8716862492263318,
|
| 539 |
+
"epoch": 2.0,
|
| 540 |
+
"frac_reward_zero_std": 0.4,
|
| 541 |
+
"grad_norm": 0.06396484375,
|
| 542 |
+
"learning_rate": 4.0100000000000006e-05,
|
| 543 |
+
"loss": 0.0028070926666259764,
|
| 544 |
+
"num_tokens": 694321.0,
|
| 545 |
+
"reward": 0.13687500059604646,
|
| 546 |
+
"reward_std": 0.33139119744300843,
|
| 547 |
+
"rewards/dbre_reward/mean": 0.13687500059604646,
|
| 548 |
+
"rewards/dbre_reward/std": 0.3313912093639374,
|
| 549 |
+
"step": 100,
|
| 550 |
+
"step_time": 27.905723336405934
|
| 551 |
+
},
|
| 552 |
+
{
|
| 553 |
+
"clip_ratio/high_max": 0.0,
|
| 554 |
+
"clip_ratio/high_mean": 0.0,
|
| 555 |
+
"clip_ratio/low_mean": 0.0,
|
| 556 |
+
"clip_ratio/low_min": 0.0,
|
| 557 |
+
"clip_ratio/region_mean": 0.0,
|
| 558 |
+
"completions/clipped_ratio": 0.95,
|
| 559 |
+
"completions/max_length": 256.0,
|
| 560 |
+
"completions/max_terminated_length": 138.6,
|
| 561 |
+
"completions/mean_length": 251.8625,
|
| 562 |
+
"completions/mean_terminated_length": 138.6,
|
| 563 |
+
"completions/min_length": 189.8,
|
| 564 |
+
"completions/min_terminated_length": 138.6,
|
| 565 |
+
"entropy": 0.9404824480414391,
|
| 566 |
+
"epoch": 2.1,
|
| 567 |
+
"frac_reward_zero_std": 0.5,
|
| 568 |
+
"grad_norm": 0.060791015625,
|
| 569 |
+
"learning_rate": 3.960000000000001e-05,
|
| 570 |
+
"loss": -0.004697377979755402,
|
| 571 |
+
"num_tokens": 728950.0,
|
| 572 |
+
"reward": 0.08519999980926514,
|
| 573 |
+
"reward_std": 0.24869290590286255,
|
| 574 |
+
"rewards/dbre_reward/mean": 0.08519999980926514,
|
| 575 |
+
"rewards/dbre_reward/std": 0.24869290590286255,
|
| 576 |
+
"step": 105,
|
| 577 |
+
"step_time": 27.779890004795742
|
| 578 |
+
},
|
| 579 |
+
{
|
| 580 |
+
"clip_ratio/high_max": 0.0,
|
| 581 |
+
"clip_ratio/high_mean": 0.0,
|
| 582 |
+
"clip_ratio/low_mean": 0.0,
|
| 583 |
+
"clip_ratio/low_min": 0.0,
|
| 584 |
+
"clip_ratio/region_mean": 0.0,
|
| 585 |
+
"completions/clipped_ratio": 0.9625,
|
| 586 |
+
"completions/max_length": 256.0,
|
| 587 |
+
"completions/max_terminated_length": 21.8,
|
| 588 |
+
"completions/mean_length": 248.425,
|
| 589 |
+
"completions/mean_terminated_length": 19.9,
|
| 590 |
+
"completions/min_length": 171.6,
|
| 591 |
+
"completions/min_terminated_length": 18.0,
|
| 592 |
+
"entropy": 1.0435175843536855,
|
| 593 |
+
"epoch": 2.2,
|
| 594 |
+
"frac_reward_zero_std": 0.3,
|
| 595 |
+
"grad_norm": 0.134765625,
|
| 596 |
+
"learning_rate": 3.91e-05,
|
| 597 |
+
"loss": -0.007500007748603821,
|
| 598 |
+
"num_tokens": 763304.0,
|
| 599 |
+
"reward": 0.13616250157356263,
|
| 600 |
+
"reward_std": 0.3243652701377869,
|
| 601 |
+
"rewards/dbre_reward/mean": 0.13616250157356263,
|
| 602 |
+
"rewards/dbre_reward/std": 0.3243652701377869,
|
| 603 |
+
"step": 110,
|
| 604 |
+
"step_time": 27.559503265400416
|
| 605 |
+
},
|
| 606 |
+
{
|
| 607 |
+
"clip_ratio/high_max": 0.0,
|
| 608 |
+
"clip_ratio/high_mean": 0.0,
|
| 609 |
+
"clip_ratio/low_mean": 0.0,
|
| 610 |
+
"clip_ratio/low_min": 0.0,
|
| 611 |
+
"clip_ratio/region_mean": 0.0,
|
| 612 |
+
"completions/clipped_ratio": 0.9875,
|
| 613 |
+
"completions/max_length": 256.0,
|
| 614 |
+
"completions/max_terminated_length": 10.0,
|
| 615 |
+
"completions/mean_length": 253.425,
|
| 616 |
+
"completions/mean_terminated_length": 10.0,
|
| 617 |
+
"completions/min_length": 214.8,
|
| 618 |
+
"completions/min_terminated_length": 10.0,
|
| 619 |
+
"entropy": 0.9914286866784096,
|
| 620 |
+
"epoch": 2.3,
|
| 621 |
+
"frac_reward_zero_std": 0.2,
|
| 622 |
+
"grad_norm": 0.0830078125,
|
| 623 |
+
"learning_rate": 3.86e-05,
|
| 624 |
+
"loss": -0.0037435129284858703,
|
| 625 |
+
"num_tokens": 798058.0,
|
| 626 |
+
"reward": 0.17202500104904175,
|
| 627 |
+
"reward_std": 0.37993268966674804,
|
| 628 |
+
"rewards/dbre_reward/mean": 0.17202500104904175,
|
| 629 |
+
"rewards/dbre_reward/std": 0.37993271350860597,
|
| 630 |
+
"step": 115,
|
| 631 |
+
"step_time": 27.579605548398103
|
| 632 |
+
},
|
| 633 |
+
{
|
| 634 |
+
"clip_ratio/high_max": 0.0,
|
| 635 |
+
"clip_ratio/high_mean": 0.0,
|
| 636 |
+
"clip_ratio/low_mean": 0.0,
|
| 637 |
+
"clip_ratio/low_min": 0.0,
|
| 638 |
+
"clip_ratio/region_mean": 0.0,
|
| 639 |
+
"completions/clipped_ratio": 0.95,
|
| 640 |
+
"completions/max_length": 256.0,
|
| 641 |
+
"completions/max_terminated_length": 128.6,
|
| 642 |
+
"completions/mean_length": 252.5,
|
| 643 |
+
"completions/mean_terminated_length": 117.6,
|
| 644 |
+
"completions/min_length": 209.0,
|
| 645 |
+
"completions/min_terminated_length": 106.6,
|
| 646 |
+
"entropy": 1.0246646717190742,
|
| 647 |
+
"epoch": 2.4,
|
| 648 |
+
"frac_reward_zero_std": 0.2,
|
| 649 |
+
"grad_norm": 0.0927734375,
|
| 650 |
+
"learning_rate": 3.8100000000000005e-05,
|
| 651 |
+
"loss": -0.004933140054345131,
|
| 652 |
+
"num_tokens": 832738.0,
|
| 653 |
+
"reward": 0.18355000019073486,
|
| 654 |
+
"reward_std": 0.36511489748954773,
|
| 655 |
+
"rewards/dbre_reward/mean": 0.18355000019073486,
|
| 656 |
+
"rewards/dbre_reward/std": 0.36511489748954773,
|
| 657 |
+
"step": 120,
|
| 658 |
+
"step_time": 27.58797151680337
|
| 659 |
+
},
|
| 660 |
+
{
|
| 661 |
+
"clip_ratio/high_max": 0.0,
|
| 662 |
+
"clip_ratio/high_mean": 0.0,
|
| 663 |
+
"clip_ratio/low_mean": 0.0,
|
| 664 |
+
"clip_ratio/low_min": 0.0,
|
| 665 |
+
"clip_ratio/region_mean": 0.0,
|
| 666 |
+
"completions/clipped_ratio": 0.9875,
|
| 667 |
+
"completions/max_length": 256.0,
|
| 668 |
+
"completions/max_terminated_length": 5.2,
|
| 669 |
+
"completions/mean_length": 253.125,
|
| 670 |
+
"completions/mean_terminated_length": 5.2,
|
| 671 |
+
"completions/min_length": 210.0,
|
| 672 |
+
"completions/min_terminated_length": 5.2,
|
| 673 |
+
"entropy": 0.9524292007088662,
|
| 674 |
+
"epoch": 2.5,
|
| 675 |
+
"frac_reward_zero_std": 0.3,
|
| 676 |
+
"grad_norm": 0.083984375,
|
| 677 |
+
"learning_rate": 3.76e-05,
|
| 678 |
+
"loss": -0.004205547273159027,
|
| 679 |
+
"num_tokens": 867468.0,
|
| 680 |
+
"reward": 0.22202499732375144,
|
| 681 |
+
"reward_std": 0.3661981552839279,
|
| 682 |
+
"rewards/dbre_reward/mean": 0.22202499732375144,
|
| 683 |
+
"rewards/dbre_reward/std": 0.3661981612443924,
|
| 684 |
+
"step": 125,
|
| 685 |
+
"step_time": 27.638399406394456
|
| 686 |
+
},
|
| 687 |
+
{
|
| 688 |
+
"clip_ratio/high_max": 0.0,
|
| 689 |
+
"clip_ratio/high_mean": 0.0,
|
| 690 |
+
"clip_ratio/low_mean": 0.0,
|
| 691 |
+
"clip_ratio/low_min": 0.0,
|
| 692 |
+
"clip_ratio/region_mean": 0.0,
|
| 693 |
+
"completions/clipped_ratio": 0.9875,
|
| 694 |
+
"completions/max_length": 256.0,
|
| 695 |
+
"completions/max_terminated_length": 34.4,
|
| 696 |
+
"completions/mean_length": 254.95,
|
| 697 |
+
"completions/mean_terminated_length": 34.4,
|
| 698 |
+
"completions/min_length": 239.2,
|
| 699 |
+
"completions/min_terminated_length": 34.4,
|
| 700 |
+
"entropy": 0.9268762037158013,
|
| 701 |
+
"epoch": 2.6,
|
| 702 |
+
"frac_reward_zero_std": 0.4,
|
| 703 |
+
"grad_norm": 0.060546875,
|
| 704 |
+
"learning_rate": 3.71e-05,
|
| 705 |
+
"loss": -0.002260996401309967,
|
| 706 |
+
"num_tokens": 902344.0,
|
| 707 |
+
"reward": 0.19577499628067016,
|
| 708 |
+
"reward_std": 0.3975414574146271,
|
| 709 |
+
"rewards/dbre_reward/mean": 0.19577499628067016,
|
| 710 |
+
"rewards/dbre_reward/std": 0.397541481256485,
|
| 711 |
+
"step": 130,
|
| 712 |
+
"step_time": 27.68072868139425
|
| 713 |
+
},
|
| 714 |
+
{
|
| 715 |
+
"clip_ratio/high_max": 0.0,
|
| 716 |
+
"clip_ratio/high_mean": 0.0,
|
| 717 |
+
"clip_ratio/low_mean": 0.0,
|
| 718 |
+
"clip_ratio/low_min": 0.0,
|
| 719 |
+
"clip_ratio/region_mean": 0.0,
|
| 720 |
+
"completions/clipped_ratio": 0.9875,
|
| 721 |
+
"completions/max_length": 256.0,
|
| 722 |
+
"completions/max_terminated_length": 26.6,
|
| 723 |
+
"completions/mean_length": 254.4625,
|
| 724 |
+
"completions/mean_terminated_length": 26.6,
|
| 725 |
+
"completions/min_length": 231.4,
|
| 726 |
+
"completions/min_terminated_length": 26.6,
|
| 727 |
+
"entropy": 0.9508431695401669,
|
| 728 |
+
"epoch": 2.7,
|
| 729 |
+
"frac_reward_zero_std": 0.1,
|
| 730 |
+
"grad_norm": 0.1025390625,
|
| 731 |
+
"learning_rate": 3.66e-05,
|
| 732 |
+
"loss": -0.004485464096069336,
|
| 733 |
+
"num_tokens": 937181.0,
|
| 734 |
+
"reward": 0.1941875010728836,
|
| 735 |
+
"reward_std": 0.39680722951889036,
|
| 736 |
+
"rewards/dbre_reward/mean": 0.1941875010728836,
|
| 737 |
+
"rewards/dbre_reward/std": 0.39680724740028384,
|
| 738 |
+
"step": 135,
|
| 739 |
+
"step_time": 27.79836461079831
|
| 740 |
+
},
|
| 741 |
+
{
|
| 742 |
+
"clip_ratio/high_max": 0.0,
|
| 743 |
+
"clip_ratio/high_mean": 0.0,
|
| 744 |
+
"clip_ratio/low_mean": 0.0,
|
| 745 |
+
"clip_ratio/low_min": 0.0,
|
| 746 |
+
"clip_ratio/region_mean": 0.0,
|
| 747 |
+
"completions/clipped_ratio": 1.0,
|
| 748 |
+
"completions/max_length": 256.0,
|
| 749 |
+
"completions/max_terminated_length": 0.0,
|
| 750 |
+
"completions/mean_length": 256.0,
|
| 751 |
+
"completions/mean_terminated_length": 0.0,
|
| 752 |
+
"completions/min_length": 256.0,
|
| 753 |
+
"completions/min_terminated_length": 0.0,
|
| 754 |
+
"entropy": 0.9166437476873398,
|
| 755 |
+
"epoch": 2.8,
|
| 756 |
+
"frac_reward_zero_std": 0.1,
|
| 757 |
+
"grad_norm": 0.078125,
|
| 758 |
+
"learning_rate": 3.61e-05,
|
| 759 |
+
"loss": 2.9802322387695314e-09,
|
| 760 |
+
"num_tokens": 972141.0,
|
| 761 |
+
"reward": 0.1720750018954277,
|
| 762 |
+
"reward_std": 0.3809880971908569,
|
| 763 |
+
"rewards/dbre_reward/mean": 0.1720750018954277,
|
| 764 |
+
"rewards/dbre_reward/std": 0.3809881091117859,
|
| 765 |
+
"step": 140,
|
| 766 |
+
"step_time": 27.89382164280105
|
| 767 |
+
},
|
| 768 |
+
{
|
| 769 |
+
"clip_ratio/high_max": 0.0,
|
| 770 |
+
"clip_ratio/high_mean": 0.0,
|
| 771 |
+
"clip_ratio/low_mean": 0.0,
|
| 772 |
+
"clip_ratio/low_min": 0.0,
|
| 773 |
+
"clip_ratio/region_mean": 0.0,
|
| 774 |
+
"completions/clipped_ratio": 0.95,
|
| 775 |
+
"completions/max_length": 256.0,
|
| 776 |
+
"completions/max_terminated_length": 117.2,
|
| 777 |
+
"completions/mean_length": 252.3125,
|
| 778 |
+
"completions/mean_terminated_length": 108.2,
|
| 779 |
+
"completions/min_length": 201.6,
|
| 780 |
+
"completions/min_terminated_length": 99.2,
|
| 781 |
+
"entropy": 1.039070624113083,
|
| 782 |
+
"epoch": 2.9,
|
| 783 |
+
"frac_reward_zero_std": 0.2,
|
| 784 |
+
"grad_norm": 0.08740234375,
|
| 785 |
+
"learning_rate": 3.56e-05,
|
| 786 |
+
"loss": 0.003438304364681244,
|
| 787 |
+
"num_tokens": 1006806.0,
|
| 788 |
+
"reward": 0.2208999961614609,
|
| 789 |
+
"reward_std": 0.401767635345459,
|
| 790 |
+
"rewards/dbre_reward/mean": 0.2208999961614609,
|
| 791 |
+
"rewards/dbre_reward/std": 0.4017676472663879,
|
| 792 |
+
"step": 145,
|
| 793 |
+
"step_time": 27.960635445202932
|
| 794 |
+
},
|
| 795 |
+
{
|
| 796 |
+
"clip_ratio/high_max": 0.0,
|
| 797 |
+
"clip_ratio/high_mean": 0.0,
|
| 798 |
+
"clip_ratio/low_mean": 0.0,
|
| 799 |
+
"clip_ratio/low_min": 0.0,
|
| 800 |
+
"clip_ratio/region_mean": 0.0,
|
| 801 |
+
"completions/clipped_ratio": 0.975,
|
| 802 |
+
"completions/max_length": 256.0,
|
| 803 |
+
"completions/max_terminated_length": 40.2,
|
| 804 |
+
"completions/mean_length": 253.0625,
|
| 805 |
+
"completions/mean_terminated_length": 27.7,
|
| 806 |
+
"completions/min_length": 220.0,
|
| 807 |
+
"completions/min_terminated_length": 15.2,
|
| 808 |
+
"entropy": 0.9523506201803684,
|
| 809 |
+
"epoch": 3.0,
|
| 810 |
+
"frac_reward_zero_std": 0.1,
|
| 811 |
+
"grad_norm": 0.0966796875,
|
| 812 |
+
"learning_rate": 3.51e-05,
|
| 813 |
+
"loss": -0.006567706167697906,
|
| 814 |
+
"num_tokens": 1041531.0,
|
| 815 |
+
"reward": 0.254237499833107,
|
| 816 |
+
"reward_std": 0.4319828271865845,
|
| 817 |
+
"rewards/dbre_reward/mean": 0.254237499833107,
|
| 818 |
+
"rewards/dbre_reward/std": 0.4319828271865845,
|
| 819 |
+
"step": 150,
|
| 820 |
+
"step_time": 33.83420682080032
|
| 821 |
+
},
|
| 822 |
+
{
|
| 823 |
+
"clip_ratio/high_max": 0.0,
|
| 824 |
+
"clip_ratio/high_mean": 0.0,
|
| 825 |
+
"clip_ratio/low_mean": 0.0,
|
| 826 |
+
"clip_ratio/low_min": 0.0,
|
| 827 |
+
"clip_ratio/region_mean": 0.0,
|
| 828 |
+
"completions/clipped_ratio": 0.95,
|
| 829 |
+
"completions/max_length": 256.0,
|
| 830 |
+
"completions/max_terminated_length": 135.8,
|
| 831 |
+
"completions/mean_length": 254.4875,
|
| 832 |
+
"completions/mean_terminated_length": 134.9,
|
| 833 |
+
"completions/min_length": 236.4,
|
| 834 |
+
"completions/min_terminated_length": 134.0,
|
| 835 |
+
"entropy": 0.9468083322048187,
|
| 836 |
+
"epoch": 3.1,
|
| 837 |
+
"frac_reward_zero_std": 0.2,
|
| 838 |
+
"grad_norm": 0.0703125,
|
| 839 |
+
"learning_rate": 3.46e-05,
|
| 840 |
+
"loss": -0.003655475750565529,
|
| 841 |
+
"num_tokens": 1076370.0,
|
| 842 |
+
"reward": 0.2434374988079071,
|
| 843 |
+
"reward_std": 0.4292252540588379,
|
| 844 |
+
"rewards/dbre_reward/mean": 0.2434374988079071,
|
| 845 |
+
"rewards/dbre_reward/std": 0.4292252600193024,
|
| 846 |
+
"step": 155,
|
| 847 |
+
"step_time": 35.7378796559904
|
| 848 |
+
},
|
| 849 |
+
{
|
| 850 |
+
"clip_ratio/high_max": 0.0,
|
| 851 |
+
"clip_ratio/high_mean": 0.0,
|
| 852 |
+
"clip_ratio/low_mean": 0.0,
|
| 853 |
+
"clip_ratio/low_min": 0.0,
|
| 854 |
+
"clip_ratio/region_mean": 0.0,
|
| 855 |
+
"completions/clipped_ratio": 0.9875,
|
| 856 |
+
"completions/max_length": 256.0,
|
| 857 |
+
"completions/max_terminated_length": 10.4,
|
| 858 |
+
"completions/mean_length": 253.45,
|
| 859 |
+
"completions/mean_terminated_length": 10.4,
|
| 860 |
+
"completions/min_length": 215.2,
|
| 861 |
+
"completions/min_terminated_length": 10.4,
|
| 862 |
+
"entropy": 0.9374438695609569,
|
| 863 |
+
"epoch": 3.2,
|
| 864 |
+
"frac_reward_zero_std": 0.3,
|
| 865 |
+
"grad_norm": 0.0693359375,
|
| 866 |
+
"learning_rate": 3.41e-05,
|
| 867 |
+
"loss": -0.005657447874546051,
|
| 868 |
+
"num_tokens": 1111126.0,
|
| 869 |
+
"reward": 0.2539374977350235,
|
| 870 |
+
"reward_std": 0.4268993496894836,
|
| 871 |
+
"rewards/dbre_reward/mean": 0.2539374977350235,
|
| 872 |
+
"rewards/dbre_reward/std": 0.4268993675708771,
|
| 873 |
+
"step": 160,
|
| 874 |
+
"step_time": 33.11068361419602
|
| 875 |
+
},
|
| 876 |
+
{
|
| 877 |
+
"clip_ratio/high_max": 0.0,
|
| 878 |
+
"clip_ratio/high_mean": 0.0,
|
| 879 |
+
"clip_ratio/low_mean": 0.0,
|
| 880 |
+
"clip_ratio/low_min": 0.0,
|
| 881 |
+
"clip_ratio/region_mean": 0.0,
|
| 882 |
+
"completions/clipped_ratio": 0.925,
|
| 883 |
+
"completions/max_length": 256.0,
|
| 884 |
+
"completions/max_terminated_length": 97.2,
|
| 885 |
+
"completions/mean_length": 244.7,
|
| 886 |
+
"completions/mean_terminated_length": 95.4,
|
| 887 |
+
"completions/min_length": 93.6,
|
| 888 |
+
"completions/min_terminated_length": 93.6,
|
| 889 |
+
"entropy": 1.0759656712412835,
|
| 890 |
+
"epoch": 3.3,
|
| 891 |
+
"frac_reward_zero_std": 0.1,
|
| 892 |
+
"grad_norm": 0.09716796875,
|
| 893 |
+
"learning_rate": 3.3600000000000004e-05,
|
| 894 |
+
"loss": 0.0013453811407089233,
|
| 895 |
+
"num_tokens": 1145182.0,
|
| 896 |
+
"reward": 0.1820499964058399,
|
| 897 |
+
"reward_std": 0.37800283133983614,
|
| 898 |
+
"rewards/dbre_reward/mean": 0.1820499964058399,
|
| 899 |
+
"rewards/dbre_reward/std": 0.37800286114215853,
|
| 900 |
+
"step": 165,
|
| 901 |
+
"step_time": 33.76342459159205
|
| 902 |
+
},
|
| 903 |
+
{
|
| 904 |
+
"clip_ratio/high_max": 0.0,
|
| 905 |
+
"clip_ratio/high_mean": 0.0,
|
| 906 |
+
"clip_ratio/low_mean": 0.0,
|
| 907 |
+
"clip_ratio/low_min": 0.0,
|
| 908 |
+
"clip_ratio/region_mean": 0.0,
|
| 909 |
+
"completions/clipped_ratio": 0.925,
|
| 910 |
+
"completions/max_length": 256.0,
|
| 911 |
+
"completions/max_terminated_length": 169.2,
|
| 912 |
+
"completions/mean_length": 250.3125,
|
| 913 |
+
"completions/mean_terminated_length": 168.7,
|
| 914 |
+
"completions/min_length": 168.2,
|
| 915 |
+
"completions/min_terminated_length": 168.2,
|
| 916 |
+
"entropy": 0.9453672260046005,
|
| 917 |
+
"epoch": 3.4,
|
| 918 |
+
"frac_reward_zero_std": 0.1,
|
| 919 |
+
"grad_norm": 0.0927734375,
|
| 920 |
+
"learning_rate": 3.3100000000000005e-05,
|
| 921 |
+
"loss": -0.001752069965004921,
|
| 922 |
+
"num_tokens": 1179687.0,
|
| 923 |
+
"reward": 0.21851249933242797,
|
| 924 |
+
"reward_std": 0.4150461137294769,
|
| 925 |
+
"rewards/dbre_reward/mean": 0.21851249933242797,
|
| 926 |
+
"rewards/dbre_reward/std": 0.4150461256504059,
|
| 927 |
+
"step": 170,
|
| 928 |
+
"step_time": 35.29977663640748
|
| 929 |
+
},
|
| 930 |
+
{
|
| 931 |
+
"clip_ratio/high_max": 0.0,
|
| 932 |
+
"clip_ratio/high_mean": 0.0,
|
| 933 |
+
"clip_ratio/low_mean": 0.0,
|
| 934 |
+
"clip_ratio/low_min": 0.0,
|
| 935 |
+
"clip_ratio/region_mean": 0.0,
|
| 936 |
+
"completions/clipped_ratio": 0.9625,
|
| 937 |
+
"completions/max_length": 256.0,
|
| 938 |
+
"completions/max_terminated_length": 82.6,
|
| 939 |
+
"completions/mean_length": 251.5625,
|
| 940 |
+
"completions/mean_terminated_length": 82.6,
|
| 941 |
+
"completions/min_length": 185.0,
|
| 942 |
+
"completions/min_terminated_length": 82.6,
|
| 943 |
+
"entropy": 0.9627169869840145,
|
| 944 |
+
"epoch": 3.5,
|
| 945 |
+
"frac_reward_zero_std": 0.0,
|
| 946 |
+
"grad_norm": 0.091796875,
|
| 947 |
+
"learning_rate": 3.26e-05,
|
| 948 |
+
"loss": -0.010714849084615707,
|
| 949 |
+
"num_tokens": 1214292.0,
|
| 950 |
+
"reward": 0.27426249384880064,
|
| 951 |
+
"reward_std": 0.42773920893669126,
|
| 952 |
+
"rewards/dbre_reward/mean": 0.27426249384880064,
|
| 953 |
+
"rewards/dbre_reward/std": 0.42773920893669126,
|
| 954 |
+
"step": 175,
|
| 955 |
+
"step_time": 32.60600872279319
|
| 956 |
+
},
|
| 957 |
+
{
|
| 958 |
+
"clip_ratio/high_max": 0.0,
|
| 959 |
+
"clip_ratio/high_mean": 0.0,
|
| 960 |
+
"clip_ratio/low_mean": 0.0,
|
| 961 |
+
"clip_ratio/low_min": 0.0,
|
| 962 |
+
"clip_ratio/region_mean": 0.0,
|
| 963 |
+
"completions/clipped_ratio": 0.9875,
|
| 964 |
+
"completions/max_length": 256.0,
|
| 965 |
+
"completions/max_terminated_length": 8.2,
|
| 966 |
+
"completions/mean_length": 253.3125,
|
| 967 |
+
"completions/mean_terminated_length": 8.2,
|
| 968 |
+
"completions/min_length": 213.0,
|
| 969 |
+
"completions/min_terminated_length": 8.2,
|
| 970 |
+
"entropy": 1.0290320612490178,
|
| 971 |
+
"epoch": 3.6,
|
| 972 |
+
"frac_reward_zero_std": 0.0,
|
| 973 |
+
"grad_norm": 0.09912109375,
|
| 974 |
+
"learning_rate": 3.21e-05,
|
| 975 |
+
"loss": -0.005978656560182571,
|
| 976 |
+
"num_tokens": 1249037.0,
|
| 977 |
+
"reward": 0.2384750008583069,
|
| 978 |
+
"reward_std": 0.4241094350814819,
|
| 979 |
+
"rewards/dbre_reward/mean": 0.2384750008583069,
|
| 980 |
+
"rewards/dbre_reward/std": 0.4241094350814819,
|
| 981 |
+
"step": 180,
|
| 982 |
+
"step_time": 33.03327611140848
|
| 983 |
+
},
|
| 984 |
+
{
|
| 985 |
+
"clip_ratio/high_max": 0.0,
|
| 986 |
+
"clip_ratio/high_mean": 0.0,
|
| 987 |
+
"clip_ratio/low_mean": 0.0,
|
| 988 |
+
"clip_ratio/low_min": 0.0,
|
| 989 |
+
"clip_ratio/region_mean": 0.0,
|
| 990 |
+
"completions/clipped_ratio": 0.925,
|
| 991 |
+
"completions/max_length": 256.0,
|
| 992 |
+
"completions/max_terminated_length": 132.8,
|
| 993 |
+
"completions/mean_length": 246.1875,
|
| 994 |
+
"completions/mean_terminated_length": 131.9,
|
| 995 |
+
"completions/min_length": 131.0,
|
| 996 |
+
"completions/min_terminated_length": 131.0,
|
| 997 |
+
"entropy": 0.92076805382967,
|
| 998 |
+
"epoch": 3.7,
|
| 999 |
+
"frac_reward_zero_std": 0.0,
|
| 1000 |
+
"grad_norm": 0.09619140625,
|
| 1001 |
+
"learning_rate": 3.16e-05,
|
| 1002 |
+
"loss": -0.009967343509197235,
|
| 1003 |
+
"num_tokens": 1283212.0,
|
| 1004 |
+
"reward": 0.33720000386238097,
|
| 1005 |
+
"reward_std": 0.46260204911231995,
|
| 1006 |
+
"rewards/dbre_reward/mean": 0.33720000386238097,
|
| 1007 |
+
"rewards/dbre_reward/std": 0.4626020550727844,
|
| 1008 |
+
"step": 185,
|
| 1009 |
+
"step_time": 34.82510126640555
|
| 1010 |
+
},
|
| 1011 |
+
{
|
| 1012 |
+
"clip_ratio/high_max": 0.0,
|
| 1013 |
+
"clip_ratio/high_mean": 0.0,
|
| 1014 |
+
"clip_ratio/low_mean": 0.0,
|
| 1015 |
+
"clip_ratio/low_min": 0.0,
|
| 1016 |
+
"clip_ratio/region_mean": 0.0,
|
| 1017 |
+
"completions/clipped_ratio": 0.95,
|
| 1018 |
+
"completions/max_length": 256.0,
|
| 1019 |
+
"completions/max_terminated_length": 74.2,
|
| 1020 |
+
"completions/mean_length": 248.875,
|
| 1021 |
+
"completions/mean_terminated_length": 45.4,
|
| 1022 |
+
"completions/min_length": 170.2,
|
| 1023 |
+
"completions/min_terminated_length": 16.6,
|
| 1024 |
+
"entropy": 1.0851470515131951,
|
| 1025 |
+
"epoch": 3.8,
|
| 1026 |
+
"frac_reward_zero_std": 0.2,
|
| 1027 |
+
"grad_norm": 0.09765625,
|
| 1028 |
+
"learning_rate": 3.1100000000000004e-05,
|
| 1029 |
+
"loss": 0.006586405634880066,
|
| 1030 |
+
"num_tokens": 1317602.0,
|
| 1031 |
+
"reward": 0.2771000027656555,
|
| 1032 |
+
"reward_std": 0.43828830122947693,
|
| 1033 |
+
"rewards/dbre_reward/mean": 0.2771000027656555,
|
| 1034 |
+
"rewards/dbre_reward/std": 0.43828831911087035,
|
| 1035 |
+
"step": 190,
|
| 1036 |
+
"step_time": 36.24654102979984
|
| 1037 |
+
},
|
| 1038 |
+
{
|
| 1039 |
+
"clip_ratio/high_max": 0.0,
|
| 1040 |
+
"clip_ratio/high_mean": 0.0,
|
| 1041 |
+
"clip_ratio/low_mean": 0.0,
|
| 1042 |
+
"clip_ratio/low_min": 0.0,
|
| 1043 |
+
"clip_ratio/region_mean": 0.0,
|
| 1044 |
+
"completions/clipped_ratio": 0.95,
|
| 1045 |
+
"completions/max_length": 256.0,
|
| 1046 |
+
"completions/max_terminated_length": 103.4,
|
| 1047 |
+
"completions/mean_length": 249.6625,
|
| 1048 |
+
"completions/mean_terminated_length": 103.4,
|
| 1049 |
+
"completions/min_length": 154.6,
|
| 1050 |
+
"completions/min_terminated_length": 103.4,
|
| 1051 |
+
"entropy": 0.9322984531521797,
|
| 1052 |
+
"epoch": 3.9,
|
| 1053 |
+
"frac_reward_zero_std": 0.0,
|
| 1054 |
+
"grad_norm": 0.099609375,
|
| 1055 |
+
"learning_rate": 3.06e-05,
|
| 1056 |
+
"loss": -0.011462598294019698,
|
| 1057 |
+
"num_tokens": 1352055.0,
|
| 1058 |
+
"reward": 0.30278749763965607,
|
| 1059 |
+
"reward_std": 0.4452593445777893,
|
| 1060 |
+
"rewards/dbre_reward/mean": 0.30278749763965607,
|
| 1061 |
+
"rewards/dbre_reward/std": 0.4452593445777893,
|
| 1062 |
+
"step": 195,
|
| 1063 |
+
"step_time": 29.027419628202914
|
| 1064 |
+
},
|
| 1065 |
+
{
|
| 1066 |
+
"clip_ratio/high_max": 0.0,
|
| 1067 |
+
"clip_ratio/high_mean": 0.0,
|
| 1068 |
+
"clip_ratio/low_mean": 0.0,
|
| 1069 |
+
"clip_ratio/low_min": 0.0,
|
| 1070 |
+
"clip_ratio/region_mean": 0.0,
|
| 1071 |
+
"completions/clipped_ratio": 0.975,
|
| 1072 |
+
"completions/max_length": 256.0,
|
| 1073 |
+
"completions/max_terminated_length": 21.8,
|
| 1074 |
+
"completions/mean_length": 250.9625,
|
| 1075 |
+
"completions/mean_terminated_length": 21.8,
|
| 1076 |
+
"completions/min_length": 175.4,
|
| 1077 |
+
"completions/min_terminated_length": 21.8,
|
| 1078 |
+
"entropy": 1.0117496035993099,
|
| 1079 |
+
"epoch": 4.0,
|
| 1080 |
+
"frac_reward_zero_std": 0.0,
|
| 1081 |
+
"grad_norm": 0.10009765625,
|
| 1082 |
+
"learning_rate": 3.01e-05,
|
| 1083 |
+
"loss": -0.017712239921092988,
|
| 1084 |
+
"num_tokens": 1386612.0,
|
| 1085 |
+
"reward": 0.2791749984025955,
|
| 1086 |
+
"reward_std": 0.43711588978767396,
|
| 1087 |
+
"rewards/dbre_reward/mean": 0.2791749984025955,
|
| 1088 |
+
"rewards/dbre_reward/std": 0.4371159017086029,
|
| 1089 |
+
"step": 200,
|
| 1090 |
+
"step_time": 27.529542883395333
|
| 1091 |
+
},
|
| 1092 |
+
{
|
| 1093 |
+
"clip_ratio/high_max": 0.0,
|
| 1094 |
+
"clip_ratio/high_mean": 0.0,
|
| 1095 |
+
"clip_ratio/low_mean": 0.0,
|
| 1096 |
+
"clip_ratio/low_min": 0.0,
|
| 1097 |
+
"clip_ratio/region_mean": 0.0,
|
| 1098 |
+
"completions/clipped_ratio": 0.9875,
|
| 1099 |
+
"completions/max_length": 256.0,
|
| 1100 |
+
"completions/max_terminated_length": 24.4,
|
| 1101 |
+
"completions/mean_length": 254.325,
|
| 1102 |
+
"completions/mean_terminated_length": 24.4,
|
| 1103 |
+
"completions/min_length": 229.2,
|
| 1104 |
+
"completions/min_terminated_length": 24.4,
|
| 1105 |
+
"entropy": 1.0013573169708252,
|
| 1106 |
+
"epoch": 4.1,
|
| 1107 |
+
"frac_reward_zero_std": 0.0,
|
| 1108 |
+
"grad_norm": 0.0908203125,
|
| 1109 |
+
"learning_rate": 2.96e-05,
|
| 1110 |
+
"loss": 0.006696997582912445,
|
| 1111 |
+
"num_tokens": 1421438.0,
|
| 1112 |
+
"reward": 0.333887505531311,
|
| 1113 |
+
"reward_std": 0.46158010959625245,
|
| 1114 |
+
"rewards/dbre_reward/mean": 0.333887505531311,
|
| 1115 |
+
"rewards/dbre_reward/std": 0.46158013343811033,
|
| 1116 |
+
"step": 205,
|
| 1117 |
+
"step_time": 27.625831649597966
|
| 1118 |
+
},
|
| 1119 |
+
{
|
| 1120 |
+
"clip_ratio/high_max": 0.0,
|
| 1121 |
+
"clip_ratio/high_mean": 0.0,
|
| 1122 |
+
"clip_ratio/low_mean": 0.0,
|
| 1123 |
+
"clip_ratio/low_min": 0.0,
|
| 1124 |
+
"clip_ratio/region_mean": 0.0,
|
| 1125 |
+
"completions/clipped_ratio": 1.0,
|
| 1126 |
+
"completions/max_length": 256.0,
|
| 1127 |
+
"completions/max_terminated_length": 0.0,
|
| 1128 |
+
"completions/mean_length": 256.0,
|
| 1129 |
+
"completions/mean_terminated_length": 0.0,
|
| 1130 |
+
"completions/min_length": 256.0,
|
| 1131 |
+
"completions/min_terminated_length": 0.0,
|
| 1132 |
+
"entropy": 1.0225385420024395,
|
| 1133 |
+
"epoch": 4.2,
|
| 1134 |
+
"frac_reward_zero_std": 0.0,
|
| 1135 |
+
"grad_norm": 0.1103515625,
|
| 1136 |
+
"learning_rate": 2.91e-05,
|
| 1137 |
+
"loss": -5.960464477539063e-09,
|
| 1138 |
+
"num_tokens": 1456398.0,
|
| 1139 |
+
"reward": 0.3235000044107437,
|
| 1140 |
+
"reward_std": 0.45962073802948,
|
| 1141 |
+
"rewards/dbre_reward/mean": 0.3235000044107437,
|
| 1142 |
+
"rewards/dbre_reward/std": 0.4596207320690155,
|
| 1143 |
+
"step": 210,
|
| 1144 |
+
"step_time": 27.678065907207202
|
| 1145 |
+
},
|
| 1146 |
+
{
|
| 1147 |
+
"clip_ratio/high_max": 0.0,
|
| 1148 |
+
"clip_ratio/high_mean": 0.0,
|
| 1149 |
+
"clip_ratio/low_mean": 0.0,
|
| 1150 |
+
"clip_ratio/low_min": 0.0,
|
| 1151 |
+
"clip_ratio/region_mean": 0.0,
|
| 1152 |
+
"completions/clipped_ratio": 0.9,
|
| 1153 |
+
"completions/max_length": 256.0,
|
| 1154 |
+
"completions/max_terminated_length": 153.2,
|
| 1155 |
+
"completions/mean_length": 244.6,
|
| 1156 |
+
"completions/mean_terminated_length": 113.6,
|
| 1157 |
+
"completions/min_length": 125.2,
|
| 1158 |
+
"completions/min_terminated_length": 74.0,
|
| 1159 |
+
"entropy": 1.0468589030206203,
|
| 1160 |
+
"epoch": 4.3,
|
| 1161 |
+
"frac_reward_zero_std": 0.3,
|
| 1162 |
+
"grad_norm": 0.07666015625,
|
| 1163 |
+
"learning_rate": 2.86e-05,
|
| 1164 |
+
"loss": -0.0048739627003669735,
|
| 1165 |
+
"num_tokens": 1490446.0,
|
| 1166 |
+
"reward": 0.22982499301433562,
|
| 1167 |
+
"reward_std": 0.4037831902503967,
|
| 1168 |
+
"rewards/dbre_reward/mean": 0.22982499301433562,
|
| 1169 |
+
"rewards/dbre_reward/std": 0.40378319621086123,
|
| 1170 |
+
"step": 215,
|
| 1171 |
+
"step_time": 27.696738472400465
|
| 1172 |
+
},
|
| 1173 |
+
{
|
| 1174 |
+
"clip_ratio/high_max": 0.0,
|
| 1175 |
+
"clip_ratio/high_mean": 0.0,
|
| 1176 |
+
"clip_ratio/low_mean": 0.0,
|
| 1177 |
+
"clip_ratio/low_min": 0.0,
|
| 1178 |
+
"clip_ratio/region_mean": 0.0,
|
| 1179 |
+
"completions/clipped_ratio": 0.9875,
|
| 1180 |
+
"completions/max_length": 256.0,
|
| 1181 |
+
"completions/max_terminated_length": 46.4,
|
| 1182 |
+
"completions/mean_length": 255.7,
|
| 1183 |
+
"completions/mean_terminated_length": 46.4,
|
| 1184 |
+
"completions/min_length": 251.2,
|
| 1185 |
+
"completions/min_terminated_length": 46.4,
|
| 1186 |
+
"entropy": 1.034738614410162,
|
| 1187 |
+
"epoch": 4.4,
|
| 1188 |
+
"frac_reward_zero_std": 0.0,
|
| 1189 |
+
"grad_norm": 0.10498046875,
|
| 1190 |
+
"learning_rate": 2.8100000000000005e-05,
|
| 1191 |
+
"loss": 0.0011565253138542176,
|
| 1192 |
+
"num_tokens": 1525382.0,
|
| 1193 |
+
"reward": 0.3366374969482422,
|
| 1194 |
+
"reward_std": 0.4732167422771454,
|
| 1195 |
+
"rewards/dbre_reward/mean": 0.3366374969482422,
|
| 1196 |
+
"rewards/dbre_reward/std": 0.47321674823760984,
|
| 1197 |
+
"step": 220,
|
| 1198 |
+
"step_time": 27.690395271993474
|
| 1199 |
+
},
|
| 1200 |
+
{
|
| 1201 |
+
"clip_ratio/high_max": 0.0,
|
| 1202 |
+
"clip_ratio/high_mean": 0.0,
|
| 1203 |
+
"clip_ratio/low_mean": 0.0,
|
| 1204 |
+
"clip_ratio/low_min": 0.0,
|
| 1205 |
+
"clip_ratio/region_mean": 0.0,
|
| 1206 |
+
"completions/clipped_ratio": 0.9875,
|
| 1207 |
+
"completions/max_length": 256.0,
|
| 1208 |
+
"completions/max_terminated_length": 27.0,
|
| 1209 |
+
"completions/mean_length": 254.4875,
|
| 1210 |
+
"completions/mean_terminated_length": 27.0,
|
| 1211 |
+
"completions/min_length": 231.8,
|
| 1212 |
+
"completions/min_terminated_length": 27.0,
|
| 1213 |
+
"entropy": 1.001887033134699,
|
| 1214 |
+
"epoch": 4.5,
|
| 1215 |
+
"frac_reward_zero_std": 0.1,
|
| 1216 |
+
"grad_norm": 0.1064453125,
|
| 1217 |
+
"learning_rate": 2.7600000000000003e-05,
|
| 1218 |
+
"loss": -0.004408703744411468,
|
| 1219 |
+
"num_tokens": 1560221.0,
|
| 1220 |
+
"reward": 0.3251750037074089,
|
| 1221 |
+
"reward_std": 0.4448351562023163,
|
| 1222 |
+
"rewards/dbre_reward/mean": 0.3251750037074089,
|
| 1223 |
+
"rewards/dbre_reward/std": 0.4448351800441742,
|
| 1224 |
+
"step": 225,
|
| 1225 |
+
"step_time": 27.640458112402122
|
| 1226 |
+
},
|
| 1227 |
+
{
|
| 1228 |
+
"clip_ratio/high_max": 0.0,
|
| 1229 |
+
"clip_ratio/high_mean": 0.0,
|
| 1230 |
+
"clip_ratio/low_mean": 0.0,
|
| 1231 |
+
"clip_ratio/low_min": 0.0,
|
| 1232 |
+
"clip_ratio/region_mean": 0.0,
|
| 1233 |
+
"completions/clipped_ratio": 0.9875,
|
| 1234 |
+
"completions/max_length": 256.0,
|
| 1235 |
+
"completions/max_terminated_length": 38.6,
|
| 1236 |
+
"completions/mean_length": 255.2125,
|
| 1237 |
+
"completions/mean_terminated_length": 38.6,
|
| 1238 |
+
"completions/min_length": 243.4,
|
| 1239 |
+
"completions/min_terminated_length": 38.6,
|
| 1240 |
+
"entropy": 0.9832980304956436,
|
| 1241 |
+
"epoch": 4.6,
|
| 1242 |
+
"frac_reward_zero_std": 0.1,
|
| 1243 |
+
"grad_norm": 0.09375,
|
| 1244 |
+
"learning_rate": 2.7100000000000005e-05,
|
| 1245 |
+
"loss": -0.0029202304780483247,
|
| 1246 |
+
"num_tokens": 1595118.0,
|
| 1247 |
+
"reward": 0.28887500166893004,
|
| 1248 |
+
"reward_std": 0.44461851716041567,
|
| 1249 |
+
"rewards/dbre_reward/mean": 0.28887500166893004,
|
| 1250 |
+
"rewards/dbre_reward/std": 0.4446185290813446,
|
| 1251 |
+
"step": 230,
|
| 1252 |
+
"step_time": 27.879659301796345
|
| 1253 |
+
},
|
| 1254 |
+
{
|
| 1255 |
+
"clip_ratio/high_max": 0.0,
|
| 1256 |
+
"clip_ratio/high_mean": 0.0,
|
| 1257 |
+
"clip_ratio/low_mean": 0.0,
|
| 1258 |
+
"clip_ratio/low_min": 0.0,
|
| 1259 |
+
"clip_ratio/region_mean": 0.0,
|
| 1260 |
+
"completions/clipped_ratio": 0.875,
|
| 1261 |
+
"completions/max_length": 256.0,
|
| 1262 |
+
"completions/max_terminated_length": 181.2,
|
| 1263 |
+
"completions/mean_length": 242.6625,
|
| 1264 |
+
"completions/mean_terminated_length": 157.15,
|
| 1265 |
+
"completions/min_length": 123.4,
|
| 1266 |
+
"completions/min_terminated_length": 123.4,
|
| 1267 |
+
"entropy": 1.0489521712064742,
|
| 1268 |
+
"epoch": 4.7,
|
| 1269 |
+
"frac_reward_zero_std": 0.1,
|
| 1270 |
+
"grad_norm": 0.10546875,
|
| 1271 |
+
"learning_rate": 2.6600000000000003e-05,
|
| 1272 |
+
"loss": -0.017240646481513976,
|
| 1273 |
+
"num_tokens": 1629011.0,
|
| 1274 |
+
"reward": 0.31321250200271605,
|
| 1275 |
+
"reward_std": 0.46374436616897585,
|
| 1276 |
+
"rewards/dbre_reward/mean": 0.31321250200271605,
|
| 1277 |
+
"rewards/dbre_reward/std": 0.46374437808990476,
|
| 1278 |
+
"step": 235,
|
| 1279 |
+
"step_time": 28.430774746800306
|
| 1280 |
+
},
|
| 1281 |
+
{
|
| 1282 |
+
"clip_ratio/high_max": 0.0,
|
| 1283 |
+
"clip_ratio/high_mean": 0.0,
|
| 1284 |
+
"clip_ratio/low_mean": 0.0,
|
| 1285 |
+
"clip_ratio/low_min": 0.0,
|
| 1286 |
+
"clip_ratio/region_mean": 0.0,
|
| 1287 |
+
"completions/clipped_ratio": 0.975,
|
| 1288 |
+
"completions/max_length": 256.0,
|
| 1289 |
+
"completions/max_terminated_length": 49.6,
|
| 1290 |
+
"completions/mean_length": 252.7,
|
| 1291 |
+
"completions/mean_terminated_length": 49.6,
|
| 1292 |
+
"completions/min_length": 203.2,
|
| 1293 |
+
"completions/min_terminated_length": 49.6,
|
| 1294 |
+
"entropy": 0.9966989070177078,
|
| 1295 |
+
"epoch": 4.8,
|
| 1296 |
+
"frac_reward_zero_std": 0.0,
|
| 1297 |
+
"grad_norm": 0.11083984375,
|
| 1298 |
+
"learning_rate": 2.61e-05,
|
| 1299 |
+
"loss": 0.0005952320992946625,
|
| 1300 |
+
"num_tokens": 1663707.0,
|
| 1301 |
+
"reward": 0.41918750405311583,
|
| 1302 |
+
"reward_std": 0.47992355227470396,
|
| 1303 |
+
"rewards/dbre_reward/mean": 0.41918750405311583,
|
| 1304 |
+
"rewards/dbre_reward/std": 0.47992355227470396,
|
| 1305 |
+
"step": 240,
|
| 1306 |
+
"step_time": 27.903805371007184
|
| 1307 |
+
},
|
| 1308 |
+
{
|
| 1309 |
+
"clip_ratio/high_max": 0.0,
|
| 1310 |
+
"clip_ratio/high_mean": 0.0,
|
| 1311 |
+
"clip_ratio/low_mean": 0.0,
|
| 1312 |
+
"clip_ratio/low_min": 0.0,
|
| 1313 |
+
"clip_ratio/region_mean": 0.0,
|
| 1314 |
+
"completions/clipped_ratio": 0.9625,
|
| 1315 |
+
"completions/max_length": 256.0,
|
| 1316 |
+
"completions/max_terminated_length": 43.8,
|
| 1317 |
+
"completions/mean_length": 250.95,
|
| 1318 |
+
"completions/mean_terminated_length": 38.3,
|
| 1319 |
+
"completions/min_length": 186.4,
|
| 1320 |
+
"completions/min_terminated_length": 32.8,
|
| 1321 |
+
"entropy": 0.9779959842562675,
|
| 1322 |
+
"epoch": 4.9,
|
| 1323 |
+
"frac_reward_zero_std": 0.0,
|
| 1324 |
+
"grad_norm": 0.083984375,
|
| 1325 |
+
"learning_rate": 2.5600000000000002e-05,
|
| 1326 |
+
"loss": -0.0041348889470100405,
|
| 1327 |
+
"num_tokens": 1698263.0,
|
| 1328 |
+
"reward": 0.3233749955892563,
|
| 1329 |
+
"reward_std": 0.4586354970932007,
|
| 1330 |
+
"rewards/dbre_reward/mean": 0.3233749955892563,
|
| 1331 |
+
"rewards/dbre_reward/std": 0.4586355030536652,
|
| 1332 |
+
"step": 245,
|
| 1333 |
+
"step_time": 27.984692039596847
|
| 1334 |
+
},
|
| 1335 |
+
{
|
| 1336 |
+
"clip_ratio/high_max": 0.0,
|
| 1337 |
+
"clip_ratio/high_mean": 0.0,
|
| 1338 |
+
"clip_ratio/low_mean": 0.0,
|
| 1339 |
+
"clip_ratio/low_min": 0.0,
|
| 1340 |
+
"clip_ratio/region_mean": 0.0,
|
| 1341 |
+
"completions/clipped_ratio": 0.9625,
|
| 1342 |
+
"completions/max_length": 256.0,
|
| 1343 |
+
"completions/max_terminated_length": 55.8,
|
| 1344 |
+
"completions/mean_length": 251.9125,
|
| 1345 |
+
"completions/mean_terminated_length": 47.0,
|
| 1346 |
+
"completions/min_length": 191.8,
|
| 1347 |
+
"completions/min_terminated_length": 38.2,
|
| 1348 |
+
"entropy": 0.9665410064160824,
|
| 1349 |
+
"epoch": 5.0,
|
| 1350 |
+
"frac_reward_zero_std": 0.1,
|
| 1351 |
+
"grad_norm": 0.10498046875,
|
| 1352 |
+
"learning_rate": 2.51e-05,
|
| 1353 |
+
"loss": 0.0026274655014276505,
|
| 1354 |
+
"num_tokens": 1732896.0,
|
| 1355 |
+
"reward": 0.37041249573230745,
|
| 1356 |
+
"reward_std": 0.46484237909317017,
|
| 1357 |
+
"rewards/dbre_reward/mean": 0.37041249573230745,
|
| 1358 |
+
"rewards/dbre_reward/std": 0.46484237909317017,
|
| 1359 |
+
"step": 250,
|
| 1360 |
+
"step_time": 27.929177726019407
|
| 1361 |
+
}
|
| 1362 |
+
],
|
| 1363 |
+
"logging_steps": 5,
|
| 1364 |
+
"max_steps": 500,
|
| 1365 |
+
"num_input_tokens_seen": 1732896,
|
| 1366 |
+
"num_train_epochs": 10,
|
| 1367 |
+
"save_steps": 50,
|
| 1368 |
+
"stateful_callbacks": {
|
| 1369 |
+
"TrainerControl": {
|
| 1370 |
+
"args": {
|
| 1371 |
+
"should_epoch_stop": false,
|
| 1372 |
+
"should_evaluate": false,
|
| 1373 |
+
"should_log": false,
|
| 1374 |
+
"should_save": true,
|
| 1375 |
+
"should_training_stop": false
|
| 1376 |
+
},
|
| 1377 |
+
"attributes": {}
|
| 1378 |
+
}
|
| 1379 |
+
},
|
| 1380 |
+
"total_flos": 0.0,
|
| 1381 |
+
"train_batch_size": 2,
|
| 1382 |
+
"trial_name": null,
|
| 1383 |
+
"trial_params": null
|
| 1384 |
+
}
|
grpo_dbre/checkpoint-250/training_args.bin
ADDED
|
Binary file (7.12 kB). View file
|
|
|