










![Paper-lightgrey)]()
Welcome to the official repository for Boogu-Image-0.1 !
English | 中文
---
> ## ⚠️ Important Notice
>
> The Boogu team does NOT currently provide any paid API, subscription, or commercial service for Boogu-Image. Any paid product or service offered under the name "Boogu-Image" — or any similar / variant name such as booguimage, Boogu Image, Boogu, etc. — is NOT affiliated with this project and is unofficial. Please verify carefully before making any payment, and stay vigilant to protect your personal privacy and financial safety.
>
> Boogu-Image-0.1 is a research project only, and not an official model release.
Boogu-Image-0.1 is a competitive Apache-2.0 open-source unified image generation and editing model family, including Base, Turbo, Edit, and other variants that provide stable, practical capabilities for high-quality text-to-image generation, fast generation, image editing, and Chinese-English text rendering. Closed-source multimodal understanding and generation systems like Nano Banana Pro and GPT-Image-2 achieve remarkable performance not because of a single model, but through a highly unified suite of system capabilities. However, under training compute that is extremely limited compared with closed-source systems, we find that systematically improving a model's understanding ability, data quality, and training pipeline can still significantly improve image generation and editing performance. Specifically, compared with some existing open-source models, our training data scale is roughly one order of magnitude smaller. We hope our empirical study and open-source release will help advance the open-source ecosystem for multimodal generation and understanding.
This repository provides checkpoints and inference code for Boogu-Image-0.1.
Since we could not evaluate on LM Arena directly, we built Boogu Arena, an LM Arena-style preference evaluation. We use an LLM to generate diverse user personas, then ask each persona to produce image generation prompts, resulting in 1K+ test prompts that we will release publicly for community reproduction. The ELO leaderboard below spans leading closed- and open-source systems. We welcome teams with questions about the results to contact us so that we can work toward a more objective, fair, and reproducible evaluation.







!Showcase for Poster & Product

> ? For the full set of practical lessons and an honest account of current limitations, see Responsible AI & Limitations below.
Beyond overall arena rankings, we break performance down by scenario across leading open-source peers. Ratings reflect our internal evaluation of typical prompts in each category.
| Model | Realistic Photography | Simple Text Rendering | Dense Text Rendering |
|---|---|---|---|
| Boogu-Image-0.1-Turbo | ⭐⭐⭐⭐ | ⭐⭐⭐⭐ | ⭐⭐⭐ |
| Boogu-Image-0.1-Base | ⭐⭐⭐ | ⭐⭐⭐⭐ | ⭐⭐⭐⭐ |
| Z-Image-Turbo | ⭐⭐⭐⭐ | ⭐⭐⭐ | ⭐⭐ |
| Qwen-Image-2512 | ⭐⭐⭐ | ⭐⭐⭐⭐ | ⭐⭐⭐ |
| Model | Params | Training | Steps | CFG | Task | Hugging Face | ModelScope | Demo |
|---|---|---|---|---|---|---|---|---|
| Boogu-Image-0.1-Base | 10B | Joint Training | 25~50 | 2.0~5.0 (e.g., 4.0) | T2I |  |  |  |
| Boogu-Image-0.1-Base-fp8 | 10B | Joint Training | 25~50 | 2.0~5.0 (e.g., 4.0) | T2I |  |  | — |
| Boogu-Image-0.1-Edit | 10B | Joint Training | 25~50 | 2.0~5.0 (e.g., 5.0) | TI2I |  |  |  |
| Boogu-Image-0.1-Edit-fp8 | 10B | Joint Training | 25~50 | 2.0~5.0 (e.g., 5.0) | TI2I |  |  | — |
| Boogu-Image-0.1-Turbo | 10B | + Decoupled DMD | 4 | 1.0 | T2I |  |  |  |
| Boogu-Image-0.1-Turbo-fp8 | 10B | + Decoupled DMD | 4 | 1.0 | T2I |  |  | — |
> Tested environment: Python 3.10 · CUDA 12.6 · PyTorch 2.7.1
# Use a brand new conda environment
conda create -y -n boogu python=3.10
conda activate boogu
# Instal necessary dependencies
# PyTorch up to 2.11.0 with CUDA up to 12.8 is supported
# Check `requirements/<torch>_<cuda>.txt`
pip install -r requirements/torch2.7-cu126.txt
pip install -e .
python utils/get_flash_attn.pyor
bash quick_start.sh
conda activate booguDownload the model weights into a local models/ directory before running inference. We recommend using the official Hugging Face CLI:
pip install -U "huggingface_hub[cli]"
# Download to ./models/<model-name>
huggingface-cli download Boogu/Boogu-Image-0.1-Base --local-dir models/Boogu-Image-0.1-Base
huggingface-cli download Boogu/Boogu-Image-0.1-Turbo --local-dir models/Boogu-Image-0.1-Turbo
huggingface-cli download Boogu/Boogu-Image-0.1-Edit --local-dir models/Boogu-Image-0.1-EditExample layout after download:
models/
└── Boogu-Image-0.1-Base/
├── model_index.json
├── mllm
├── processor
├── scheduler
├── transformer
└── vaeThen point inference to the local path via --model models/Boogu-Image-0.1-Base.
This repository provides utils/get_flash_attn.py to automatically install a compatible flash-attn wheel for your environment.
Requirements:
# Auto: detect environment, download a prebuilt wheel, fallback to source build
python utils/get_flash_attn.py
# Force source compilation
python utils/get_flash_attn.py --buildThe script first searches <code>mjun0812/flash-attention-prebuild-wheels</code>, then tries official <code>Dao-AILab/flash-attention</code> release wheels with both cxx11abi variants, and finally falls back to source compilation via pip install flash-attn --no-build-isolation.
export device="cuda:0" # Required
mkdir -p outputs/test_turbo/
# Prompt enhancement is powered by an instruction reasoner, also called the rewriter.
# We provide two ways to use it:
#
# 1. Standalone external rewriter:
# See utils/t2i_external_prompt_rewriter.py. This is a pure external mode example and
# requires enough GPU memory, without advanced memory management.
# python utils/t2i_external_prompt_rewriter.py --prompt "draw a cat" --model /path/to/Qwen3-VL-32B-Instruct --lang en
#
# 2. Pipeline-integrated rewriter:
# See the scripts under `demo_scripts` whose names contain "reasoning".
# For example: demo_scripts/demo_t2i_local_reasoning.sh
# This mode supports more flexible memory management. Set the generation and
# rewriter devices manually, then pass them to inference.py:
# export device="cuda:0"
# export rewriter_device="cuda:1"
# python inference.py --device $device --rewriter_device $rewriter_device ...
# For more details, see INFERENCE_GUIDE.md.
python inference_turbo.py
--pretrained_pipeline_name_or_path "models/Boogu-Image-0.1-Turbo"
--instruction "一幅国风琉金风格的山水画作,展现了桂林山水在金光普照下的壮丽景象。远山层叠,江水如镜,山峰边缘勾勒着发光的金色线条。画面采用石青石绿岩彩与鎏金质感相结合,局部有厚涂油画笔触,空中飘浮着金色粒子,营造出梦幻朦胧而又磅礴大气的意境。"
--height 2048 --width 2048
--output_image_path "outputs/test_turbo/out_1.png"
--device "$device"> ? For full CLI options, device setup, offload strategies, caching acceleration, Torch Compile, FP8, and batch inference details, see <strong>INFERENCE_GUIDE.md</strong>.
> Torch Compile note: --enable_torch_compile can occasionally produce all-black outputs on some GPUs/models. If that happens, disable it first.
| VRAM | Recommended Config (T2I 1K) | Recommended Config (T2I 2K) |
|---|---|---|
| 12GB | Unquantized: --enable_sequential_cpu_offload_flag Quantized: --enable_model_cpu_offload_flag --use_fp8_weights | Unquantized: --enable_sequential_cpu_offload_flag Quantized: --enable_group_offload_flag --use_fp8_weights |
| 16GB | Unquantized: --enable_sequential_cpu_offload_flag Quantized: --enable_model_cpu_offload_flag --use_fp8_weights | Unquantized: --enable_sequential_cpu_offload_flag Quantized: --enable_model_cpu_offload_flag --use_fp8_weights |
| 24GB | Unquantized: --enable_model_cpu_offload_flag Quantized --use_fp8_weights | --enable_model_cpu_offload_flag |
| 32GB | Unquantized: --enable_model_cpu_offload_flag Quantized: --use_fp8_weights | Unquantized: --enable_model_cpu_offload_flag Quantized: --use_fp8_weights |
| 40GB | Base Model | Unquantized: --enable_model_cpu_offload_flag Quantized: --use_fp8_weights |
| 80GB | Base Model | Base Model |
Boogu-Image-0.1 is released for research purposes and is not intended for production deployment without additional safeguards. We took responsible-AI considerations into account during data curation, training, and evaluation; however the model may still produce outputs that are inaccurate, biased, or otherwise inappropriate.
? World Knowledge Gap
?️ Image-to-Image Consistency & In-Context Scenarios
? Text Rendering Stability
? Body Structure in Complex Poses
? Small Faces & Small Limbs
? Limited Release Scope
Downstream users are responsible for applying content moderation, validation, and compliance checks appropriate to their use case.
Closed-source systems such as GPT-Image, Nano Banana, and the Seedream series helped us understand the frontier capabilities and practical boundaries of unified understanding-and-generation systems. We thank the Qwen-Image, Z-Image, OmniGen2, FLUX, and broader open-source communities for the foundations they provide, and DeepSeek for strong open-source understanding models that support open-source unified multimodal systems.
This project is released under the Apache-2.0 License.
本文按 GitCode 项目 README 搬运。README front matter 标注 license: apache-2.0。页面公开来源与授权字段已保留。
来源:https://gitcode.com/hf_mirrors/Boogu/Boogu-Image-0.1-Turbo