Amazon JumpStart Text Generation
Introduction to SageMaker Built-In Algorithms - Text Generation
This notebook's CI test result for us-west-2 is as follows. CI test results in other regions can be found at the end of the notebook.
Welcome to Amazon SageMaker Built-In Algorithms! You can use Sagemaker Built-In Algorithms to solve many Machine Learning tasks through one-click in SageMaker Studio, or through SageMaker Python SDK.
In this demo notebook, we demonstrate how to use the SageMaker Python SDK for Text Generation. Text generation is the task of generating text which appears indistinguishable from the human-written text. It is also sometimes known as "natural language generation". Here, we show how to use state-of-the-art pre-trained GPT models for Text Generation. We also demonstrate running inference on any Text Generation model available on HugginFace
Note: This notebook was tested on ml.t3.medium instance in Amazon SageMaker Studio with Python 3 (Data Science) kernel and in Amazon SageMaker Notebook instance with conda_pytorch_p39 kernel.
1. Set Up
Before executing the notebook, there are some initial steps required for set up. This notebook requires ipywidgets.
Permissions and environment variables
To host on Amazon SageMaker, we need to set up and authenticate the use of AWS services. Here, we use the execution role associated with the current notebook as the AWS account role with SageMaker access.
2. Select a pre-trained model
You can continue with the default model, or can choose a different model from the dropdown generated upon running the next cell. A complete list of SageMaker pre-trained models can also be accessed at JumpStart pre-trained Models.
[Optional] Select a different Sagemaker pre-trained model. Here, we download the model_manifest file from the Built-In Algorithms s3 bucket, filter-out all the Text Generation models and select a model for inference.
Chose a model for Inference
Using Models not Present in the Dropdown
If you want to choose any other model which is not present in the dropdown and is available at HugginFace Text Generation please choose huggingface-textgeneration-models in the dropdown and pass the model_id in the HF_MODEL_ID variable. Inference on the models listed in the dropdown menu can be run in network isolation. In such a case, no inbound or outbound network calls can be made to or from the model container. The models listed in the dropdown can also be deployed with custom VPC settings, which can provide your model container with a network connection within your VPC that is not connected to the internet. Refer to AWS documentation for more details.
However, when running inference on a model specified through HF_MODEL_ID, the model container will download the model artifact from HuggingFace. Therefore, the model container cannot run in network isolation. Furthermore, if you want to use custom VPC settings, you must provide access the HuggingFace portal in your VPC.
3. Retrieve Artifacts & Deploy an Endpoint
Using SageMaker, we can perform inference on the pre-trained model, even without fine-tuning it first on a new dataset. We start by retrieving the deploy_image_uri, deploy_source_uri, and model_uri for the pre-trained model. To host the pre-trained model, we create an instance of sagemaker.model.Model and deploy it. This may take a few minutes.
4. Query endpoint and parse response
Input to the endpoint is any string of text dumped in json and encoded in utf-8 format. Output of the endpoint is a json with generated text.
Below, we put in some example input text. You can put in any text and the model predicts next words in the sequence. Longer sequences of text can be generated by calling the model repeatedly.
5. Advanced features
This model also supports many advanced parameters while performing inference. They include:
- max_length: Model generates text until the output length (which includes the input context length) reaches
max_length. If specified, it must be a positive integer. - num_return_sequences: Number of output sequences returned. If specified, it must be a positive integer.
- num_beams: Number of beams used in the greedy search. If specified, it must be integer greater than or equal to
num_return_sequences. - no_repeat_ngram_size: Model ensures that a sequence of words of
no_repeat_ngram_sizeis not repeated in the output sequence. If specified, it must be a positive integer greater than 1. - temperature: Controls the randomness in the output. Higher temperature results in output sequence with low-probability words and lower temperature results in output sequence with high-probability words. If
temperature-> 0, it results in greedy decoding. If specified, it must be a positive float. - early_stopping: If True, text generation is finished when all beam hypotheses reach the end of stence token. If specified, it must be boolean.
- do_sample: If True, sample the next word as per the likelyhood. If specified, it must be boolean.
- top_k: In each step of text generation, sample from only the
top_kmost likely words. If specified, it must be a positive integer. - top_p: In each step of text generation, sample from the smallest possible set of words with cumulative probability
top_p. If specified, it must be a float between 0 and 1. - seed: Fix the randomized state for reproducibility. If specified, it must be an integer.
- return_full_text: If True, input text will be part of the output generated text. If specified, it must be boolean. The default value for it is False.
We may specify any subset of the parameters mentioned above while invoking an endpoint. Next, we show an example of how to invoke endpoint with these arguments
6. Clean up the endpoint
Notebook CI Test Results
This notebook was tested in multiple regions. The test results are as follows, except for us-west-2 which is shown at the top of the notebook.