Haojie Zheng1, 2, Yixin Yang2 , Siqi Yang2, Shuchen Weng 1,2* , Boxin Shi 2*
* Corresponding authors
1 BAAI, 2 Peking University
InstructAV2AV.mp4 |
| For the best experience, please enable audio. |
We are working hard to deliver the components as soon as possible. Here is our roadmap for the upcoming release:
- Inference Code (Including pre-trained weights and pipeline scripts)
- InsAVE Dataset (The complete dataset used for training and evaluation)
- Training Scripts (Codes for reproducing or fine-tuning our model)
For CUDA 12.1, you can install the dependencies with the following commands.
git clone https://github.com/suimuc/InstructAV2AV.git
cd InstructAV2AV
conda create -n instructav2av python=3.10 -y
conda activate instructav2av
pip install torch==2.6.0 torchvision torchaudio
pip install -r requirements.txt
pip install flash_attn --no-build-isolation
pip install -e .Using huggingface-cli to download the models:
# InstructAV2AV Model
huggingface-cli download suimu/InstructAV2AV \
--local-dir ckpts/InstructAV2AV
# Text Encoder & Video VAE
huggingface-cli download Wan-AI/Wan2.2-TI2V-5B \
--include \
"google/*" \
"models_t5_umt5-xxl-enc-bf16.pth" \
"Wan2.2_VAE.pth" \
--local-dir ckpts/Wan2.2-TI2V-5B
# Audio VAE
huggingface-cli download hkchengrex/MMAudio \
--include \
"ext_weights/best_netG.pt" \
"ext_weights/v1-16.pth" \
--local-dir ckpts/MMAudioThe unified inference entry point is scripts/edit.py. The source video must contain an audio track unless a separate file is provided with --source-audio.
python scripts/edit.py \
--source-video assets/input.mp4 \
--instruction "Keep the person’s identity and change the spoken words to <S>This is more than just art, it’s a statement.<E>." \
--finetune-path ckpts/InstructAV2AV/clone_id_voice.safetensors \
--output outputs/edited.mp4To achieve optimal editing quality and better text-to-video alignment, you can enhance the instruction by appending a detailed description of the final edited video immediately after your editing command.
python scripts/edit.py \
--source-video assets/input1.mp4 \
--instruction "Change the man into a young woman with brown hair, wearing a gray blazer over a light pink top and a necklace with a heart-shaped pendant, and saying, <S>I really think we should give it another chance.<E>. A young woman with long, wavy brown hair and fair skin is engaged in a conversation. She is wearing a gray blazer over a light pink top and has a necklace with a heart-shaped pendant. Her facial expressions change throughout the sequence, showing a range of emotions that suggest she is either explaining something earnestly or reacting to a conversation, and says <S>I really think we should give it another chance.<E> The setting appears to be indoors, with a dimly lit, blurred background that suggests a social environment, possibly a bar or restaurant. The focus remains on the woman's face, capturing her reactions and engagement in the dialogue." \
--finetune-path ckpts/InstructAV2AV/general.safetensors \
--output outputs/edited.mp4InstructAV2AV provides multiple task-specific checkpoints. All checkpoints use the same model architecture, but each checkpoint is fine-tuned for a different editing type. Select the checkpoint that best matches your instruction and provide its path through --finetune-path.
| Checkpoint | Editing Type | Description | Example Instruction |
|---|---|---|---|
general |
General Edit | Flexible editing of appearance, scenes, actions, speech, and sound. | Make the horse dark brown with a white saddle. |
insertion |
Content Insertion | Add an object or other content to the source video. | Add a dark vintage sedan driving from the right to the left. |
removal |
Content Removal | Remove an object and its associated audiovisual content from the source video. | Remove the chipmunk standing on the stone surface among the peanuts. |
clone_id |
Identity Cloning | Preserve a person's visual identity while editing other visual or audio attributes. | Keep the person's appearance, change the timbre to a man, and change the spoken words to <S>I understand, but I think we need to consider.<E>. |
clone_voice |
Voice Cloning | Preserve the speaker's timbre while editing the video or spoken content. | Keep the timbre, change the person's appearance, and change the spoken words to <S>I came here to tell you that you should go.<E>. |
clone_id_voice |
Identity + Voice Cloning | Preserve both the person's visual identity and the speaker's timbre. | Keep the person's identity and voice, and change the spoken words to <S>This is more than just art, it's a statement.<E>. |
After downloading all six InstructAV2AV checkpoints and the dependency weights, launch the demo with:
python scripts/demo.py --shareThe training dataset is available at suimu/InsAVE-80K.
The trainer accepts CSV, JSON, or JSONL manifests with the following canonical fields:
| Field | Description |
|---|---|
source_video |
Original video |
source_audio |
Original audio |
target_video |
Edited video |
target_audio |
Edited audio |
instruction |
Editing instruction |
The unified training entry point is scripts/train.py, with configuration in ovi/configs/train/train_av_edit.yaml.
Before training, update at least these entries:
ckpt_dir: ./ckpts
finetune_path: ./ckpts/InstructAV2AV/general.safetensors
dataset:
metadata_path: ./data/InsAVE-80K/path/to/manifest.csvThe manifest and initialization checkpoint can also be overridden from the command line:
accelerate launch \
--config_file ovi/configs/train/accelerate_config.yaml \
scripts/train.py \
--config-file ovi/configs/train/train_av_edit.yaml \
--data-manifest data/InsAVE-80K/path/to/manifest.csv \
--finetune-path ckpts/InstructAV2AV/general.safetensors \
--output-dir outputs/train_av_editThe provided Accelerate configuration uses DeepSpeed ZeRO-2 and is configured for eight processes. Change num_processes in ovi/configs/train/accelerate_config.yaml to match the number of available GPUs.
The training objective is selected automatically from the modality flags:
has_video |
has_audio |
Training objective |
|---|---|---|
true |
true |
Joint audio-video editing |
true |
false |
Video-only editing |
false |
true |
Audio-only editing |
This project builds on the following open-source projects:
- Ovi for the audio-video fusion backbone.
- Wan2.2 for the video backbone components, UMT5 text encoder, and video VAE.
- MMAudio for the audio VAE and vocoder.
We thank the authors and contributors of these projects for releasing their work.
If you find our work or dataset helpful for your research, please consider citing our paper:
@article{instructav2av2026,
title={InstructAV2AV: Instruction-Guided Audio-Video Joint Editing},
author={Zheng, Haojie and Yang, Yixin and Yang, Siqi and Weng, Shuchen and Shi, Boxin},
journal={arXiv preprint arXiv:2605.18467},
year={2026}
}