michaelfeil
/

ct2fast-codegen-350M-mono

michaelfeil commited on May 21, 2023

Commit

a6788a4

1 Parent(s): a3834ca

Upload Salesforce/codegen-350M-mono ctranslate fp16 weights

Files changed (1) hide show

README.md CHANGED Viewed

@@ -11,7 +11,7 @@ Speedup inference while reducing memory by 2x-4x using int8 inference in C++ on
 quantized version of [Salesforce/codegen-350M-mono](https://huggingface.co/Salesforce/codegen-350M-mono)
 ```bash
-pip install hf-hub-ctranslate2>=2.0.7
 ```
 Converted on 2023-05-21 using
 ```
@@ -33,10 +33,11 @@ model = GeneratorCT2fromHfHub(
         model_name_or_path=model_name,
         device="cuda",
         compute_type="int8_float16",
-        tokenizer=AutoTokenizer.from_pretrained("Salesforce/codegen-350M-mono")
 )
 outputs = model.generate(
     text=["def print_hello_world():", "def hello_name(name:"],
 )
 print(outputs)
 ```

 quantized version of [Salesforce/codegen-350M-mono](https://huggingface.co/Salesforce/codegen-350M-mono)
 ```bash
+pip install hf-hub-ctranslate2>=2.0.8
 ```
 Converted on 2023-05-21 using
 ```
         model_name_or_path=model_name,
         device="cuda",
         compute_type="int8_float16",
+        # tokenizer=AutoTokenizer.from_pretrained("Salesforce/codegen-350M-mono")
 )
 outputs = model.generate(
     text=["def print_hello_world():", "def hello_name(name:"],
+    max_length=64
 )
 print(outputs)
 ```