Instructions to use lokeshrao226/eco-logistics-qwen-grpo with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use lokeshrao226/eco-logistics-qwen-grpo with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("unsloth/qwen2.5-1.5b-instruct-unsloth-bnb-4bit") model = PeftModel.from_pretrained(base_model, "lokeshrao226/eco-logistics-qwen-grpo") - Notebooks
- Google Colab
- Kaggle
Eco-Logistics GRPO Agent
LoRA adapter for Qwen-2.5-1.5B-Instruct, post-trained with GRPO on the Eco-Logistics OpenEnv environment β a 3-warehouse supply chain task that forces the agent to balance profit against carbon emissions under non-stationary demand shocks.
Submitted to the OpenEnv Hackathon β India 2026 (theme: World Modeling for Professional Tasks).
Headline result
On the held-out hardest task (net_zero_profit), 3-run averaged across 30 episodes:
25.6Γ improvement in profit-per-carbon ratio vs the base model Grader 0.273 Β± 0.019 (vs base Qwen 0.259)
The trained policy generalizes from the medium training task (inventory_balanced) to the harder evaluation task (net_zero_profit) with stable variance across re-runs.
Visual comparison
Cross-task evaluation
We evaluate the same trained model on two tasks. Both use held-out seeds 500β509.
Task 1 β inventory_balanced (training-distribution, single 10-episode run)
| Policy | Profit | Carbon | Profit/Carbon | Grader |
|---|---|---|---|---|
| Random | 3946 | 1252 | 3.2 | 0.065 |
| Heuristic | 3558 | 450 | 7.9 | 0.040 |
| Base Qwen-2.5-1.5B | 3793 | 1429 | 2.7 | 0.135 |
| GRPO Qwen (this model) | 4828 | 122 | 39.6 | 0.195 |
Task 2 β net_zero_profit (held-out, 3-run averaged, 30 episodes)
| Policy | Profit | Carbon | Profit/Carbon | Grader |
|---|---|---|---|---|
| Random | 2636.6 | 1076.8 | 2.85 | 0.001 |
| Heuristic | 3735.2 | 0.0 | β | 0.292 |
| Base Qwen-2.5-1.5B | 3708.8 | 25.2 | 2.65 | 0.259 |
| GRPO Qwen (this model) | 3687.9 Β± 55.7 | 54.3 Β± 64.9 | 67.96 | 0.273 Β± 0.019 |
We trust the net_zero_profit numbers more β 3-run averaging gives Ο=0.019 on grader, and inventory_balanced showed substantial single-run variance (re-runs gave profit/carbon of 4421/695 and 4767/217). The heuristic still wins on grader (0.292), but our learned policy closes ~67% of the gap between base Qwen and the heuristic without any env-specific rules.
Usage
from unsloth import FastLanguageModel
from peft import PeftModel
import torch
model, tokenizer = FastLanguageModel.from_pretrained(
model_name = "Qwen/Qwen2.5-1.5B-Instruct",
max_seq_length = 2048,
dtype = None,
load_in_4bit = True,
)
model = PeftModel.from_pretrained(model, "lokeshrao226/eco-logistics-qwen-grpo")
FastLanguageModel.for_inference(model)
The model expects a ChatML prompt with the system message describing the 3-warehouse task and the initial observation. It outputs a JSON array of 10 actions (one per episode step). Full prompt template and rollout loop are in the training notebook.
Training details
- Method: GRPO via TRL + Unsloth, LoRA r=16, 4-bit quantization
- Hardware: Single T4 GPU (Colab)
- Steps: 30 GRPO steps, learning rate 2e-6
- Dataset: 50 unique initial-state prompts across all 3 tasks, seeds 0β49
- Held-out eval: seeds 500β509 (never seen during training)
- Design choice β Upfront Trajectory Planning: the model emits the entire 10-step plan as one JSON array per inference. Cuts HTTP round-trips by 10Γ, makes GRPO-over-HTTP tractable on a T4.
Anti-reward-hacking patch (critical detail)
The first training run collapsed at step 15: the model discovered that outputting invalid JSON triggered a SAFE_FALLBACK_ACTIONS plan that shipped almost nothing, which meant near-zero carbon penalty. The β5.0 format penalty was smaller in magnitude than the carbon savings, so the model learned to output garbage on purpose. Fix: change the format penalty to β1000.0, making invalid output strictly worse than any reachable valid reward. Full debugging story in the writeup.
A documented negative result: SFT-then-GRPO
We attempted to fix our weak valid-action rate (~20%) with an SFT warmup phase. SFT did fix valid-rate (20% β 80%), but subsequent GRPO collapsed in 3 different reward configurations. Diagnosis: SFT made the policy too uniform, so GRPO had no variance among the N=4 sampled completions to learn from. Detailed in the writeup. We submit the no-SFT model as our headline; the SFT experiment is documented as honest evidence of what we tried.
Limitations
- Valid-action rate of 20% on held-out seeds. The biggest weakness β model produces parseable JSON only 1 in 5 times; the rest fall through to a heuristic fallback.
- Heuristic still beats us on grader (0.292 vs 0.273). Not a SOTA claim.
- Upfront planning constraint β agent commits to all 10 steps at t=0 without intermediate replanning.
- Profit didn't generalize from
inventory_balanced(+27%) tonet_zero_profit(essentially flat). The carbon-efficiency story is what travels across tasks.
Links
- Code & writeup: https://github.com/Lokeshrao69/eco-logistics
- Live environment: https://huggingface.co/spaces/lokeshrao226/eco-logistics
- Training notebook:
train_eco_logistics_grpo_v8_FINAL.ipynb
Team: Crystal Blue. Built in 3 days for OpenEnv Hackathon (India 2026).
- Downloads last month
- 8

