Skip to main content

Command Palette

Search for a command to run...

Decoding AI Jargons

Understanding how AI model works

Published
•5 min read•View as Markdown
Decoding AI Jargons

In the world of Artificial Intelligence (AI), almost 8 out 10 developer’s are starting their generative AI juronery, and I am one of them. Before starting with Gen AI, you should be aware of some important AI concepts. In this article we going to understand some core concepts of AI.

Everyone has used Chat GPT, but have you wondered how the GPT answers your query?, how it writes a code that fits to your requirements?, How it generates the content?, Let’s understand this.

Well, the answer of this is in the name of Chat GPT - GPT (Generative Pretrained Transformer), it is a Transformer which is Pertained with high volume of data and it is generative in nature. Let’s understander this transformer in detail.

Chat GPT is interface using which we can chat with GPT model with added functionality.

Transformer

Ref: https://research.google/pubs/attention-is-all-you-need/

The transformer you see in the above image is proposed by Google, It was originally developed for Google translator, This also used in LLM training. We won’t going the depth of transform, but only going to understand how it works

Phase 1 : Input Embedding

Tokenization

When enter a query to GPT, your query is first tokenized based on some predertimed algorithms. Every LLM has its own tokenizer.

Tokenization is nothing but breaking down the text or input data into smaller chunks called tokens. This tokens are often converted into numerical for better understating of LLMs. Tokenization is first step in making data understandable for AI

tiktoken is library used by OpenAI to convert words into tokens,

import tiktoken

encoder = tiktoken.encoding_for_model('gpt-4o')
print("Vocab Size: ", encoder.n_vocab) # (200,019)

text="The cat sat on the mat"

encoded = encoder.encode(text)
print("Encoded Tokens: ", encoded) # [976, 9059, 10139, 402, 290, 2450]

decoded = encoder.decode(encoded)
print("Decoded Text: ", decoded) # "The cat sat on the mat"

Vector Embedding

Once the your input query is tokenized, this data is transformed into multidimensional representation where each dimension represent the similar meaning and close relation among the words. This numerical repreparation of data points representing their meaning and relation is known as Vector Embedding. This enables LLM’s to process data efficiently.

The position of data points in space reflects its meaning and relation with other data points. The words with similar meaning will be placed closer in the space while words with dissimilar meaning will apart from each other.

Example:

King and Queen are similar words and has relation in between them, also the words Man and Woman has same meaning and same relation between them, when we plot this words on 3D space, it will looks something like below.

On the graph there will be same distance between King to Queen and Man to Woman, King to Man and Queen to Woman.

https://projector.tensorflow.org/ You can go to this link to visualize how vector embeddings looks like.

Vector Embeddings helps us find the semantic meaning of the words.

from google import genai
from dotenv import load_dotenv
import os

load_dotenv()

client = genai.Client(api_key=os.getenv("GEMINI_API_KEY"))

text="The cat sat on the mat"

result = client.models.embed_content(
  model="gemini-embedding-001",
  contents=text
)

print("Embeded Text: ", result.embeddings)

#  values=[
#     -0.022850024,
#     0.012087576,
#     -0.0012127935,
#     -0.07871191,
#     0.0022645553,
#     <... 3067 more items ...>,
#   ]

Phase 2 : Positional Encoding

Once the vector embeddings are generated then this embeddings are encoded based on their position. In This process the words in sentences are provided with the information about their sequence in the sentence. As the transformer process the all the input tokens parallely, and don’t capture the sequence of the words. This encoding is added to the word embeddings to allow the model to understand the relationships between words based on their position in the sentence.

Without positional information, the model would treat all words the same regardless of their position, making it difficult to understand context and relationships. Positional encoding involves adding a vector, representing the position of a token in the sequence, to the token's embedding vector.

Example

Consider the sentence “The Cat sat on the mat”, if we don’t consider the positional encoding sentence might become as “The mat sat on the Cat”, which will entirely change the meaning of the sentence.

Phase 3 : Self Attention

Self Attention is the mechanism that allows the transformer to highlight the importance of the different words of input sentence when processing the words. It helps the model to understand the relation between the words by considering the entire context rather than just relaying on static embeddings. Self Attention allows the model to dynamically adjust the representation of each word based on its relationship with all other words in the sequence, leading to a more context-aware understanding of the input.

Example:

Consider we have two sentences -

  1. The river bank

  2. The ICICI bank

In this above sentences bank is common word, but it has different meaning in both the sentences. In Self Attention it has been made aware that in first sentence it is associated with river and in second sentence it is associated with ICICI, so the word back adjust its embedding to be more precise about which back are we talking about.

It can done in single headed or multi headed format.

Both Encoder and Decoder of LLM works on the nearly same principle.

Extra key concepts

Knowledge cutoff

All the LLM’s are trained with past data and can not answer the query related to real-time data. Every model has a Knowledge cutoff, the data till the data model is trained on.

When you ask the GPT what is current weather of Pune, It won’t be able to answer it, as it does not have current data.

But Chat GPT can answer it, as it uses Agentic AI workflows to call Weather API and gets the real-time weather information. This is one of the added functionality

Temperature

It is a parameter of model that controls the randomness and creativity of the model's output

Softmax

The softmax is function that is used in the final layer of a neural network model for classification of tasks, it converts raw output scores into probabilities. it also known as logits

Vocab Size:

Number of unique token available in the tokenization dictionary