Decoding AI Jargons
In this articles we will learn how an LLM works , what are the steps it undergoes after taking the inputs and serving the output and also simplifying the jargon’s in artificial intelligence world.
GPT
G- Generative - as it generates the content according to our input (text/image/video)
P - Pretrained - as it generates on the data with which it has been trained (text/ image/video)
T- Transformer - This denotes the underlying architecture how this the process is done between input and output
Flow of LLM
- First the when the input is provided to LLM it will tokenize it.
Tokenize (encoding) → This is a process of converting the input to an array of number which are equivalent for that particular model/LLM
import tiktoken
encoder = tiktoken.encoding_for_model('gpt-4o')
print('vocab size',encoder.n_vocab)
text= "The cat sat on the mat"
tokens= encoder.encode(text)
print(tokens)
rawtext= encoder.decode(tokens)
print(rawtext)
This tokens will be a array of numbers and in the context of LLM every numbers has a meaning to it.
You can try to create tokens according to your model using this link https://tiktokenizer.vercel.app/
Vocab size → Vocab size for an LLM is dictionary for LLM. It has that limit of dictionary to convert the input to tokens. Lets say a person know hindi and english his vocab size will be more than a person who know only english.
- After tokenization the LLM will try to find meaning for the tokens through vector embeddings.
Vector embedding are addresses for the token to relate them to real world meaning called as (sematic meaning) with this the LLM can make relationships with the input data given by moving in distance and direction as per the input.
Here in the below figure we could see with the help of dog and puppy cat could get its sematic meaning with distance and direction.

After vector embedding the LLM will do the positional encoding this will give addition context to the vector embedding position because if the position of the token changes the meaning itself changes
Then the LLM does the self attention which means the embeddings will communicate with each other with single head / multi head attention( attention gets the why what how etc aspects ).
There are two things in LLM working 1. Training phase 2. Inferencing phase
Training phase is the time where we input the data and get some random output from LLM and we will take the loss function for actual output and expected ouput and feed the input again by changing the weights of the neural network until we get good outputs
Inferencing phase is the time where we use the LLM and it gives most probably good output
Softmax - is the last process where the LLM gets the most desired outputs for an input with probablities
Temperature - This can be changed , with change this might select most desired output or likely desired output from the softmax outputs
knowledge cutoff - Lets say a person studied till 7th class and we ask him about 8th syllabus he has knowledge cutoff of 7th class only, In same way the LLM will have till date knowledge cutoff