Describe the bug
We conducted tests on OPT/GPTJ/GPT-Neox/BLOOM 7B INT8, these models are all producing garbage outputs on DeepSpeed 0.8.1
OPT model is NCCL communication issue
GPT-NeoX 20B is producing garbage
BLOOM-7B: shape '[1, 4, 32, 384]' is invalid for input of size 16384
How we tested?
We generated int8 checkpoints of the model and then loaded them back. Example of doing the same with DS inference test suite.
deepspeed --num_nodes 1 \
--num_gpus 8 \
inference-test.py \
--use_kernel \
--ds_inference \
--use_meta_tensor \
--name EleutherAI/gpt-neox-20b \
--checkpoint_path /tmp/ws/gpt-neox-20b/ \
--save_mp_checkpoint_path /tmp/ws/sharded-gpt-neox-20b/ \
--dtype int8
deepspeed --num_nodes 1 \
--num_gpus 8 \
inference-test.py \
--use_kernel \
--ds_inference \
--use_meta_tensor \
--name EleutherAI/gpt-neox-20b \
--checkpoint_path /tmp/ws/sharded-gpt-neox-20b/ \
--dtype int8
More info this.
#2770
Creating a new issue to track the int8 checkpoint loading issue.
Describe the bug
We conducted tests on OPT/GPTJ/GPT-Neox/BLOOM 7B INT8, these models are all producing garbage outputs on DeepSpeed 0.8.1
How we tested?
We generated int8 checkpoints of the model and then loaded them back. Example of doing the same with DS inference test suite.
More info this.
#2770
Creating a new issue to track the int8 checkpoint loading issue.