Podcast

Every article and concept, narrated. Listen on the page, or subscribe and take them anywhere.

Subscribe to the feed Paste this feed into Apple Podcasts, Spotify, Overcast, or any podcast app.

Articles

169 narrated articles, newest first.

Finding Bard Inside Google's Gemma 4

8 August 2026 · 1:15

A mechanistic interpretability study shows how ablating specific MLP neurons in Gemma 4 causes the model to revert to its older, dormant Bard identity.

Mechanistic Interpretability of Google's Gemma - Safety and Personality Insights

6 August 2026 · 1:24

Modifying specific self-attention layers in the Gemma language model alters its self-identity and safety responses, blurring real-world harm with video games.

XM on text diffusion: a small real gain, and a bigger effect we were not looking for

4 August 2026 · 1:22

An evaluation of Explorative Modeling on a text diffusion model: a small consistent gain at matched steps, and a larger likelihood-versus-sampling tradeoff that turned out to belong to the training schedule.

Readers spot AI writing 64% of the time: the first 1,636 answers from the AI vs Human test

31 July 2026 · 1:20

First results from the dejan.ai AI vs Human test: 1,636 answers, 64% correct overall, 18 perfect rounds against 0.14 expected from guessing, and accuracy rising from 55% to 75% with reading time.

Advanced New Tail Keyword Research

27 July 2026 · 1:24

Show 100 people a photo of your product and ask what they would type into Google. A survey on Zoysia Tenuifolia grass returned 68 terms: 24 already on the top-ranking page and 44 it had never used.

Gemini picks whichever web page it reads first 92% of the time. Only pages rewritten to its measured preferences ever beat that.

26 July 2026 · 1:16

Shown two web pages, Gemini picks the first one 92% of the time. Pages rewritten to 31 measured text preferences win 15% from the second slot. Ordinary pages: 0%.

2037. AGI. China probably.

23 July 2026 · 1:36

An exploration of DeepSeek founder Liang Wenfeng's strategy of deliberate restraint, continual learning, and open weights in the race toward AGI.

No, AI doesn't prefer Reddit. Search does.

21 July 2026 · 1:20

Analysis of citation mining data shows Reddit has a 99.39% rejection rate from OpenAI, zero citations from Anthropic, and heavy usage by Google.

Engadget: Quantitative Linguistic Analysis

19 July 2026 · 1:40

This paper analyzes Engadget's article corpus from 2006 to 2023 using Shannon entropy and data compression to track editorial trends and article length.

The Verge: Quantitative Linguistic Analysis

19 July 2026 · 1:43

This paper profiles the recovered corpus of The Verge from 2011 to 2023 using Shannon entropy and data compression to analyze changes in article structure.

How Brave's AI Search Works

13 July 2026 · 1:37

DEJAN AI reveals the system instructions for Brave Ask, detailing how the AI grounding and generation pipeline processes web evidence to produce answers.

The Paradox of Average and AI Visibility

12 July 2026 · 1:39

A study by DEJAN AI examines how AI rankers prioritize prototypical content over brand authority, suggesting a dual strategy for SEO and human conversion.

Introducing eCommerce Optimization Engine

9 July 2026 · 1:32

The eCommerce Optimization Engine uses Google's multimodal embedding model to align product text with images to improve search retrieval and SEO performance.

Content Optimization Engine Insights by DEJAN AI

8 July 2026 · 1:41

We spent 2.71 Billion tokens teaching a machine to reverse-engineer how AI search ranks pages. What it found overturns the usual SEO playbook: structure and intent beat every credential you've been told to add, and the engine that proved it reached #1 in 1,554 of 2,246 runs.

Grounding Source ≠ Citation ≠ Mention

7 July 2026 · 1:42

Being a cited grounding source does not mean your brand was mentioned in the answer, and being mentioned does not mean you were cited. They are correlated but distinct events produced at different stages of the AI answer pipeline — and conflating them is the source of most AI-visibility measurement

Cascading Hallucination

7 July 2026 · 1:46

A novel failure mode in multi-agent AI systems: a capable model trusts its sub-agents to deliver grounded facts, but they hallucinate, and the orchestrator launders the fiction into a confident answer without ever questioning it.

Introducing LinkjeBERT a Dutch Language Model

5 July 2026 · 1:35

LinkjeBERT is a Dutch language model trained on structured Markdown to predict where humans will place links in text using token-level confidence levels.

Google Uses Chrome to Supply Context to Gemini Chat

5 July 2026 · 1:29

An analysis of Chrome's Page Content Agent, which builds a structured tree of page elements to provide Gemini with grounded context for reading and interaction.

Generative Self-Retrieval: How AI Models Recall Brand Facts From Memory

27 June 2026

When an AI answers about your brand from memory, generative self-retrieval decides whether it recalls you correctly or invents a plausible wrong answer.

Primary Bias

21 June 2026 · 1:35

Primary bias is what an AI model already believes about your brand before it searches: an ungrounded confidence baked into training that becomes the biggest factor in whether your content is selected in AI answers.

Grounding Snippets

21 June 2026 · 1:25

Grounding snippets are the short, extractive sentences AI systems pull from your page to build their answers: the atomic unit of AI-search visibility, where most of your page never makes the cut.

Selection Rate Optimization

21 June 2026 · 1:24

Selection Rate Optimization is the AI-search counterpart to click-through-rate optimization: improving how often AI systems choose your content as the source they ground and cite their answers on.

Relevance Engineering

21 June 2026 · 1:26

Relevance Engineering is the practice of building a page's semantic relevance to a query with embeddings and vector math, treating search visibility as an engineering problem rather than keyword optimization.

AIO (AI Optimization)

21 June 2026 · 1:00

AIO (AI Optimization) is shaping your content and presence so AI systems favor them in generated responses — one route to AI Visibility.

What is AI SEO?

21 June 2026 · 1:23

AI SEO optimizes content and brands for inclusion in generated answers. Explore how AI answers work, compare AI and traditional SEO, and view 2026 data.

GEO (Generative Engine Optimization)

21 June 2026 · 0:54

GEO (Generative Engine Optimization) is optimizing for LLM-powered search and chat so your brand and pages appear in generated answers.

AEO (Answer Engine Optimization)

21 June 2026 · 1:01

AEO (Answer Engine Optimization) is optimizing to be the answer that AI assistants and answer engines give to a question.

The Open Knowledge Format (OKF)

21 June 2026 · 1:17

Google Cloud's Open Knowledge Format is an open, vendor-neutral way to package the context AI systems need, as plain markdown files any model or agent can read.

AI Visibility

21 June 2026 · 1:16

AI Visibility is the outcome of AI-centric SEO, focusing on being cited and recommended by AI systems through brand mentions and source citations in answers.

Teaching a Model to Reason Before It Learns to Talk

19 June 2026

An exploration of building tiny, logic-first models using cellular automata to challenge the transformer paradigm and identify the primitives of reasoning.

How Search Grounding Biased an LLM Against YouTube

18 June 2026

An analysis of how Claude's webinar platform recommendations were influenced by affiliate-driven content, and a correction regarding YouTube's live features.

How AI Search Grounding Actually Works: Google vs OpenAI vs Anthropic

13 June 2026

An analysis of how Google, OpenAI, and Anthropic handle web grounding, comparing their search processes, citation rates, and how they process page content.

Emotion Geometry of Google’s AI Models

17 May 2026

A replication study of Anthropic’s emotion research on Google’s Gemma 4 31B model, finding that internal emotion representations organize along a valence axis.

Google’s (still) doesn’t see your live page.

7 May 2026

Australian AI SEO agency specialising in brand visibility optimisation for global brands and e-commerce websites using machine learning and data science.

Gemma 4 Brand Authority Map

4 April 2026

A comparison of brand recall between Google's Gemma 4 and Gemini 3 Flash models, analyzing how open-weight and closed models prioritize different brands.

Chrome’s New Shopping Classifier

3 April 2026

An analysis of Google's shopping classifier model in Chrome, detailing its content extraction pipeline, chunking logic, and impact on e-commerce SEO.

AI Brand Authority Index: Ranking 2.9 Million Brands by Associative Embeddedness in Gemini’s Memory

28 March 2026

This research presents a methodology for quantifying brand authority in large language model memory using Personalized PageRank and directed association graphs.

TurboQuant: From Paper to Triton Kernel in One Session

25 March 2026

An implementation and technical analysis of Google's TurboQuant algorithm, testing KV cache compression on Gemma 3 4B using PyTorch and custom Triton kernels.

Clickbait Titles Exploit Attention Through Latent Entities

22 March 2026

Clickbait titles function by withholding a latent entity—the subject, reason, process, or outcome—to force a click and resolve an artificial information gap.

Fanout Query Analysis

20 March 2026

An analysis of 365,920 fanout queries from Google, OpenAI, and Amazon reveals how different AI models generate internal search queries for web grounding.

Reverse Prompting: Reconstructing Prompts from AI-Generated Text

18 March 2026

A fine-tuned Gemma 3 270M model reconstructs the most likely prompts from AI-generated responses using synthetic data and contrastive search configurations.

Rufus – Under the Hood. What Drives Amazon’s AI Shopping Assistant?

15 March 2026

An overview of the technical architecture behind Amazon's Rufus, covering its query planning, RAG-based retrieval, custom LLM models, and streaming response.

Is Query Length a Reliable Predictor of Search Volume?

12 March 2026

An analysis of 39.6 million Amazon search queries reveals that query length is an unreliable predictor of search volume compared to semantic content.

Search Grounding is Transient

6 March 2026

Google’s AI search and Gemini use a single-turn transient architecture that purges raw web snippets from working memory immediately after a response is sent.

SRO & Grounding Snippets

1 March 2026

Selection Rate Optimization (SRO) is a new discipline focused on visibility in AI-powered search by measuring how often content is selected for grounding.

What extraction method is Google using to build grounding snippets?

24 February 2026

An analysis of Google's Gemini grounding pipeline, examining how extractive summarization selects query-focused sentences to build grounding context from web sources.

Implicit Queries in AI Search

24 February 2026

An analysis of Google patent US11769017B1, detailing a system that uses context and implied input engines to proactively generate and push AI summaries.

Sorry Google, I was wrong.

18 February 2026

An analysis of a $2,000 Gemini API bill caused by the URL Context tool, which ingests entire web pages as input tokens without providing size estimates.

AI Search Has a Spam Problem

18 February 2026

This article examines GEO spam, a method of manipulating AI-generated answers through self-referential content and engineered claims designed for grounding.

WebMCP

10 February 2026

WebMCP is a proposed web standard that allows websites to expose structured tools to AI agents via declarative and imperative APIs for better reliability.

Bias and Prejudice in AI Search

30 December 2025

An exploration of primary bias in AI, defined as a model's inherent confidence in an entity based on training data, and its impact on brand selection rates.

Most People Don’t Read

30 December 2025

A qualitative study comparing self-reported reading habits against actual user behavior, tracking mouse movements, scroll patterns, and time on page.

Google’s Trajectory: 2026 and Beyond

25 December 2025

Google's shift toward agentic AI involves Gemini robotics, A2UI for secure interfaces, and the AP2 protocol for autonomous agent payments and commerce.

Google’s Ranking Signals

24 December 2025

Overview of search ranking factors including popularity signals, PCTR models, semantic relevance, keyword matching, freshness, and various search modes.

How big are Google’s grounding chunks?

20 December 2025

Analysis of how Google selects content to ground Gemini-powered AI shows a fixed 2,000-word budget per query, where relevance rank determines word share.

Google’s AI Uses Schema?

20 December 2025

An investigation into whether Google uses structured data to ground Gemini in AI search, exploring the relationship between LD+JSON and RAG grounding sources.

Dynamic Visual Layouts

18 December 2025

Dynamic visual layout (DVL) is a generative user interface where layouts are created on demand to suit specific queries, shifting the focus from SEO to information.

Grounding Snippet Extraction Tool

15 December 2025

The Gemini Grounding Tool identifies which URLs and specific sentences Google's AI extracts to ground its answers, helping optimize content for AI search.

How Long Are Web Pages?

14 December 2025

An analysis of 44,684 web pages reveals a median content length of 3,201 tokens and an average of 10,403 tokens, highlighting implications for AI systems.

Google AI Search Update: Completely New Grounding Format

13 December 2025

An observation of a new, custom grounding context format for Gemini that deviates from the traditional index-based model used in previous prompt types.

AI Mode, Content & Search Index

13 December 2025

Tests suggest Google’s AI Mode uses a proprietary content store rather than retrieving live web content from the search index during the query fan out process.

How user prompts shape your content visibility in AI search.

13 December 2025

An analysis of how AI search rankers use semantic alignment to surface different content zones within a single article based on query specificity and intent.

Report: How People Use AI at Work

10 December 2025

An analysis of qualitative interviews with 1,250 professionals exploring how the general workforce, creatives, and scientists integrate AI into their work.

How do people use AI assistants?

5 December 2025

An analysis of 3.9 million AI chat sessions reveals that most interactions are short, non-commercial, and involve users seeking help with writing, learning, or coding.

Ricursive: The Most Interesting AI Company You Haven’t Heard Of

3 December 2025

Ricursive Intelligence, founded by Anna Goldie and Azalia Mirhoseini, aims to automate chip design using AI to enable recursive self-improvement in hardware.

Better Vector Clustering With Head Noun Extraction

28 November 2025

An exploration of how standard embeddings can create a semantic soup by grouping search queries by adjectives rather than head nouns during clustering.

Advanced Prompting Techniques for AI SEO

27 November 2025

Explore prompt engineering techniques for SEO, including zero-shot, few-shot, role, and chain-of-thought prompting to improve content and automate tasks.

To block or not to block? Bot is the question.

26 November 2025

An overview of AI bots, distinguishing between training data scrapers used for LLM development and agentic bots designed for autonomous, goal-oriented tasks.

Gemini 3 hallucinates fan-out queries

22 November 2025

An analysis of Gemini 3 API responses reveals the model fabricating search queries to justify its answers, demonstrating persistent hallucination behaviors.

AI SEO Deep Dive – Tom Critchlow & Dan Petrovic

19 November 2025

A deep-dive conversation with Tom Critchlow on the mechanics of AI search, focusing on Selection Rate Optimization (SRO) and how to influence LLM behavior.

OpenAI’s Sparse Circuits Breakthrough and What It Means for AI SEO

14 November 2025

OpenAI research on sparse circuits shows AI models can be built with fewer connections, making them more interpretable and easier to analyze for AI SEO.

How GPT Sees the Web

14 November 2025

A technical walkthrough of how GPT handles web search, including snippets, expansions, context size settings, and the sliding window mechanism for retrieval.

BlockRank: A Faster, Smarter Way to Rank Documents with LLMs

10 November 2025

BlockRank is a novel method for in-context ranking that uses structured sparse attention and contrastive training to improve LLM efficiency and accuracy.

In AI SEO #10 is the new #1

9 November 2025

An empirical study analyzing how Google's AI Mode uses text snippets from multiple sources, finding that snippets are more prompt-aligned than full web pages.

How much of your content survives the AI Search filter?

8 November 2025

An analysis of the Google grounding process, detailing how user prompts and source snippets are processed by models and measuring citation coverage rates.

Browsing vs Content Fetcher

8 November 2025

Google's AI Mode uses browsing for single URL retrieval and content_fetcher for batch processing of multiple structured sources within a workflow.

From Free-Text to Likert Distributions: A Practical Guide to SSR for Purchase Intent

15 October 2025

Semantic Similarity Rating (SSR) maps LLM free-text responses to Likert distributions to improve purchase intent realism and match human response patterns.

Claude System Internals

9 October 2025

An exploration of the internal processes of Claude, including system prompts, token budgets, search grounding algorithms, and hidden reasoning blocks.

CAPS: A Content Attribution Payment Scheme for the AI Era

30 September 2025

The collapse of the web's economic model due to AI is addressed through the Content Attribution Payment Scheme, a framework for micropayments and grounding.

AI Search Citation Mining

27 September 2025

Raw data dump from a citation mining pipeline demo featuring 60 prompts across AEO, AI marketing, AI optimization, AI SEO, and AIO using GPT-5 and Gemini.

Using GPT-5 Structured Output Markers to Detect AI-Generated Content Online

27 September 2025

Publishing unedited AI-generated text can leak internal GPT-5 structured output markers like turn0search21, which can lead to SEO and reputational risks.

TimesFM-ICF

26 September 2025

Google Research's TimesFM-ICF uses in-context fine-tuning to achieve high-performance time-series forecasting without the need for traditional model training.

Chrome Screen AI Protos

23 September 2025

A directory of protocol buffer files covering various machine intelligence technologies, including OCR, vision, face detection, and image classification.

RexBERT

23 September 2025

RexBERT is a domain-specialized language model trained on e-commerce text to optimize product titles, descriptions, attribute extraction, and semantic search.

Annotated Page Content (APC)

22 September 2025

Annotated Page Content (APC) is a structured protobuf representation of a webpage's layout and content, designed for actionable and efficient downstream use.

Deconstructing DomDistiller: How Chrome’s Reader Mode Algorithm Impacts Technical SEO

22 September 2025

An analysis of Chrome's DomDistiller engine explains how it uses heuristics, DOM traversal, and semantic HTML to isolate main content from page boilerplate.

LLM is a Presentation Layer in AI Search

21 September 2025

Large language models act as a presentation layer on top of classic information retrieval. They rely on crawling, indexing, and ranking to prevent hallucinations.

Gemini App Tools – A Technical Overview

14 September 2025

Gemini acts as an orchestration layer that manages a large language model by deconstructing prompts into tasks for tools like Code Interpreter and APIs.

EmbeddingGemma: The Game-Changing Model Every SEO Professional Needs to Know

5 September 2025

Google's EmbeddingGemma is a multilingual embedding model that mirrors Gemini's architecture to provide insights into semantic search and query intent.

Primary Bias on Selection Rate in AI Search

4 September 2025

Selection Rate measures how often AI systems select specific items from grounding results. It explores primary bias, model relevance, and the Tree Walker algo.

The Latent History of AI Boom

1 September 2025

An exploration of how the transition from RNNs to transformers and the discovery of double descent enabled the scaling of large language models like GPT.

AI Overviews = Dialogflow Agent?

31 August 2025

An analysis of AI Overview leaks suggesting that Google's implementation may be based on the Dialogflow agentic framework, specifically regarding intent priority.

Fan-Out Query Search Volume Prediction Using Deep Learning

30 August 2025

A deep learning approach using a Query Demand Estimator to automatically predict search volume ranges for long-tail queries generated by a fan-out model.

Comprehensive Guide to Identifying AI Comment Bots

28 August 2025

Identify AI-generated comments through statistical analysis of sentiment, formulaic linguistic patterns, repetitive vocabulary, and a lack of human imperfection.

What is “Help Me Write” in Chrome?

27 August 2025

Help Me Write is Google Chrome's AI-powered assistant that generates context-aware text suggestions for short-form content like emails, posts, and forms.

Introducing Tree Walker

24 August 2025

Tree Walker is an analysis tool designed to deconstruct how AI models like Gemini perceive brands by uncovering word uncertainty and probabilistic language paths.

Does Schema Help With “AI”?

23 August 2025

An experiment testing whether OpenAI's browsing tool provides GPT-5 with grounding context from page schema or only extracts plain text and markdown content.

Your website is about to start talking. Are you ready for this?

21 August 2025

Explore how Chrome's built-in Gemini Nano model uses semantic HTML and the accessibility tree to enable private, on-device AI conversations on websites.

Inside Chrome’s Semantic Engine: A Technical Analysis of History Embeddings

21 August 2025

Technical analysis of Chrome's history embeddings system, detailing the DocumentChunker algorithm, passage extraction, and the 1540-dimensional vector pipeline.

What does an SEO do in the AI age?

19 August 2025 · 1:20

Modern search engines use a hybrid structure consisting of a strategic Agentic Layer for decision-making and an Interpretative Layer for generative synthesis.

Understanding and Control

17 August 2025

AI optimization relies on mechanistic interpretability to understand internal neural computations and model steering to actively control model behavior.

People call them AI. That’s it.

16 August 2025

Social media poll results from 864 votes show that while AI is the dominant label for tools like ChatGPT and Claude, users remain divided on preferred terms.

GPT-5 Made SEO Irreplaceable

10 August 2025

OpenAI is shifting its model design to prioritize reasoning and intelligence over memorized world knowledge, relying on tools and retrieval for information.

Google’s Query Fan-Out System – A Technical Overview

9 August 2025

This article describes a system that replicates Google's query fan-out approach by using generative neural networks to automatically create intelligent search variants.

GPT-5 System Prompt

8 August 2025

Here it is: Credit to: https://x.com/elder_plinius/status/1953583554287562823H/T https://x.com/DarwinSantosNYC for spotting it.

Journalism Is Dead. Say Hello to Gournalism.

6 August 2025

Explores the rise of Gournalism, a shift toward generative, AI-produced content optimized for machine consumption and algorithmic indexing.

Human Friendly Content is AI Friendly Content

21 July 2025

Explore the parallels between human and AI attention mechanisms and learn how to optimize content for both through scannable structures and hierarchy.

Analysis of Gemini Embed Task-Based Dimensionality Deltas

16 July 2025

An analysis of Gemini Embed optimization modes, including classification, retrieval, and semantic similarity, through vector embedding dimension visualization.

Dynamic per-label thresholds for large-scale search query classification with Otsu’s method

9 July 2025

Explore how to use Otsu's algorithm to solve the problem of inconsistent confidence thresholds in search-query intent classifiers using dynamic, per-label tuning.

Prompt Engineer’s Guide to Gemini Schemas

2 July 2025

A technical guide to the Gemini API GenerateContentResponse schema, detailing the structure of candidates, usage metadata, safety ratings, and parsed data.

Top 10 Most Recent Papers by MUVERA Authors

30 June 2025

A collection of recent research papers and focus areas for MUVERA authors Laxman Dhulipala, Majid Hadian, Jason Lee, and Rajesh Jayaram.

Training Gemma‑3‑1B Embedding Model with LoRA

28 June 2025

Gemma-Embed is a bespoke 256-dim embedding model created by fine-tuning google/gemma-3-1b-pt with LoRA to enable high-fidelity query reformulation.

Training a Query Fan-Out Model

24 June 2025

Google generates high-quality query reformulations by traversing the mathematical latent space between queries and documents to train the qsT5 model.

Cosine Similarity or Dot Product?

19 June 2025

An examination of the Chrome codebase reveals that the history_embeddings component uses the dot product of normalized vectors to perform similarity searches.

Universal Query Classifier

13 June 2025

A zero-shot, multi-label search query classifier that maps queries to any user-provided label taxonomy without the need for retraining or bespoke models.

Another failed attempt to kill SEO

9 June 2025

An analysis of the term Generative Engine Optimization (GEO) and a critique of industry rebranding efforts following opinions shared by Andreessen Horowitz personnel.

Vector Embedding Optimization

6 June 2025

An evaluation of four embedding methods comparing speed, storage, and accuracy. Results show mrl truncation maintains high accuracy while reducing file size.

Dissecting Gemini’s Tokenizer and Token Scores

5 June 2025

Explore how Google’s Gemini processes text using subword tokenization. Use this tool to inspect SentencePiece log-likelihood scores for common and rare tokens.

There’s a small army of on-device models coming to Chrome

5 June 2025

Technical interpretations and parameter breakdowns for various AI models, including Gemini, Gemma, ULM, and StableLM, covering architecture and scale.

AI Mode Site Search

4 June 2025

Explore Vertex AI website search features, including Enterprise edition tools like extractive answers, image search, and advanced LLM capabilities for summaries.

Multi-Step Research Agent

4 June 2025

An implementation of Google's query fan-out in an agentic framework used to research the machine learning and SEO services offered by DEJAN Marketing.

Query Fan-Out Prompt Implementation in Google’s Open-Source Agentic Framework

4 June 2025

Google’s Gemini Fullstack LangGraph Quickstart uses Gemini 2.5 and LangGraph to build a citation-driven research agent with a React and FastAPI architecture.

From Hallucinations to Clicks

2 June 2025

An automated method for mapping LLM-hallucinated URLs to valid pages using keyword matching and semantic similarity via vector embeddings and cosine similarity.

What is GEO?

2 June 2025

Generative Engine Optimisation (GEO) is a term used to describe SEO for AI assistants and generative search engines, often based on a single research paper.

AI Mode & Page Indexing

30 May 2025

Tests indicate Google's AI Mode uses a proprietary content store rather than the live web, as it fails to fetch indexed pages that are otherwise ranking.

AI Mode is Not Live Web

29 May 2025

An experiment testing Google's AI Mode suggests it may rely on Google's existing index or cached web data rather than performing live HTTP requests for all URLs.

How AI Mode Selects Snippets

28 May 2025

An analysis of how Google selects content for AI Mode snippets, identifying patterns in value propositions, HTML structure, and semantic selection criteria.

AI Mode Internals

28 May 2025

An exploration of Google's AI Mode and Gemini tools, including its use of Google Search, Python libraries, and how it processes date, time, and location data.

The Future of Google

28 May 2025

Sundar Pichai discusses Google's AI strategy, the evolution of Search, upcoming AR glasses, the impact of AI on web traffic, and the future of robotics.

The Inner Workings of GPT’s file_search Tool

27 May 2025

The file_search tool allows GPT models to extract precise information from uploaded documents using structured queries and provides citations for verification.

Live Blog: Hacking Gemini Embeddings

24 May 2025

An experimental study reproducing the vec2vec research paper by attempting to translate and align Gemini and MxbAI embedding spaces using unsupervised methods.

Google’s New URL Context Tool

21 May 2025

Google's Gemini now uses a combination of search and browsing tools to fetch and read specific web pages, allowing it to ground responses in real-world data.

LLM-Based Search Volume Prediction

19 May 2025

An analysis comparing Google Gemini's keyword volume predictions against actual Google Search Console data reveals weak-to-moderate correlation and limited accuracy.

How Google grounds its LLM, Gemini.

8 May 2025

An analysis of Gemini's internal grounding processes, revealing its structured indexing method, operational stages, and use of external verification tools.

Google Lens Modes

8 May 2025

The lns_mode parameter classifies Google Lens queries into text, unimodal, or multimodal modes to help route requests and support AI Mode functionality.

Content Substance Classification

23 April 2025

Cyberfluff is a novel approach for detecting low-quality web content using curriculum-driven contrastive pretraining to distinguish fluff from substance.

Chrome’s New Embedding Model: Smaller, Faster, Same Quality

19 April 2025

Chrome's latest update features a new text embedding model that is 57% smaller than its predecessor, using int8 quantization to maintain search quality.

AI Content Detection

17 April 2025

DEJAN-LM is an AI content detection model trained on 20 million sentences, using a combined deep learning and heuristic approach to identify advanced AI text.

I think Google got it wrong with “Generate → Ground” approach.

17 April 2025

An analysis of Google's RARR framework compared to retrieval-first approaches like RAG and FiD, focusing on reducing LLM hallucinations through grounding.

Introducing Grounding Classifier

2 April 2025

An analysis of Gemini 2.5 Pro's search grounding capabilities and the development of a prompt grounding classifier trained on 10,000 collected prompts.

Advanced Interpretability Techniques for Tracing LLM Activations

31 March 2025

This page explores mechanistic interpretability techniques, including activation logging, causal tracing through activation patching, and attention head analysis.

Temperature Parameter for Controlling AI Randomness

30 March 2025

The temperature parameter in generative AI models influences randomness and creativity by rescaling the probability distribution of potential next words.

Probability Threshold for Top-p (Nucleus) Sampling

30 March 2025

Top-p sampling, or nucleus sampling, is a parameter used in generative AI to control text randomness by selecting words based on a cumulative probability.

How Google Decides When to Use Gemini Grounding for User Queries

29 March 2025

Google uses dynamic retrieval to decide when Gemini models should use grounding. A prediction score and configurable threshold determine if a query needs search data.

Cross-Model Circuit Analysis: Gemini vs. Gemma Comparison Framework

29 March 2025

A framework for comparative circuit analysis between Google's Gemini and Gemma models to identify how different architectures represent brand information.

Neural Circuit Analysis Framework for Brand Mention Optimization

29 March 2025

This framework uses open-weight models like Gemma 3 Instruct to perform mechanistic brand positioning through direct neural circuit and activation analysis.

Strategic Brand Positioning in LLMs: A Methodological Framework for Prompt Engineering and Model Behavior Analysis

29 March 2025

This paper presents a methodological framework for analyzing and optimizing brand mentions in large language models through systematic prompt probing and analysis.

AlexNet: The Deep Learning Breakthrough That Reshaped Google’s AI Strategy

21 March 2025

Google and the Computer History Museum open-sourced the AlexNet code, highlighting its role in launching deep learning and shaping Google's AI-first strategy.

The Next Chapter of Search: Get Ready to Influence the Robots

19 March 2025

Explore the evolving landscape of SEO, focusing on how AI, conversational search, and Large Language Models are changing brand representation and visibility.

Revealed: The exact search result data sent to Google’s AI.

14 March 2025

An analysis of Gemini's grounding capabilities, addressing issues with hallucinations, guardrails, and the discovery of multi-passage snippet context.

Beyond Rank Tracking: Analyzing Brand Perceptions Through Language Model Association Networks

27 February 2025

The DEJAN methodology uses large language models to analyze brand perception and semantic associations, moving beyond traditional keyword rank tracking.

Teaching AI Models to Be Better Search Engines: A New Approach to Training Data

13 February 2025

A recent patent application describes a method for training AI models to better understand human queries by using LLMs to automatically generate training data.

Self-Supervised Quantized Representation for KG-LLM Integration

6 February 2025 · 1:24

Self-Supervised Quantized Representation (SSQR) integrates knowledge graphs with large language models by compressing entity information into discrete codes.

What does Gemini think about your brand?

29 January 2025

Chrome Dev includes a quantized Gemini model for tasks like scam prevention. This analysis examines its on-device execution and reverse-engineered prompts.

Google’s Privacy Sandbox: Navigating the Cookieless Future

14 January 2025

An examination of Google's Privacy Sandbox, focusing on the technical details and privacy implications of the Topics API and the FLEDGE API.

Why deep learning works.

26 December 2024

An excerpt from François Chollet’s Deep Learning with Python exploring the manifold hypothesis and how structured information enables deep learning to work.

Introducing VecZip: Embedding Compression Algorithm

12 December 2024

VecZip is a novel compression method by DEJAN AI that reduces embedding dimensionality by retaining unique dimensions to improve AI performance and storage.

Site Engagement Metrics

29 November 2024

The Google Site Engagement Metrics Framework in Chromium tracks user interactions, engagement scores, and browsing behavior using UMA histograms.

Beyond Links: Understanding Page Transitions in Chrome

27 November 2024

Explore Chrome page transition types and qualifiers to understand user intent, navigation pathways, and the SEO implications of different browser behaviors.

Both humans and AI return similar results when asked for a random number

13 November 2024

A comparison of 200,000 random numbers provided by humans and Google's Gemma-2-2b-it model reveals significant overlaps and patterns in number selection.

Chrome AI Frameworks & Models

30 October 2024

A comprehensive list of Chrome's on-device machine learning models, including specialized tools for language processing, page analysis, and content safety.

Attention Is All You Need

13 October 2024 · 1:19

A discussion of the Attention Is All You Need paper, covering the Transformer architecture, multi-head attention, and its impact on machine translation.

The State of AI

13 October 2024 · 1:35

The 2024 State of AI report explores the rise of open models, benchmarking challenges, neurosymbolic systems, model efficiency, and global AI developments.

ILO

5 October 2024 · 1:09

The ILO app is a Streamlit-based tool for managing SEO data through URL population, GSC data fetching, query intent classification, and traffic projections.

Resource-Efficient Binary Vector Embeddings With Matryoshka Representation Learning

5 September 2024

An analysis of reducing vector embedding storage through Matryoshka Representation Learning and binary embeddings to optimize SEO text feature extraction.

Query Intent via Retrieval Augmentation and Model Distillation

5 September 2024

QUILL enhances query intent classification by using retrieval augmentation and a two-stage distillation process to balance model performance and efficiency.

Search Query Quality Classifier

31 August 2024

A search query classifier using ALBERT architecture to identify well-formed queries with 80% accuracy, improving upon Google's LSTM-based model by 10%.

How Gemini Selects Results

26 August 2024

An explanation of how internal algorithms use relevance scoring, recency bias, user intent, and stochasticity to retrieve and present information.

Gemini System Prompt

26 August 2024

Gemini Advanced provides access to the Gemini 1.5 Pro model, featuring a 1 million token context window for analyzing up to 1500 pages of information.

Concept explainers

176 narrated glossary concepts, A–Z.

A2UI (Agent-to-User Interface)

6 July 2026 · 1:11

An open web standard that lets AI agents generate rich, interactive interfaces as declarative data, rendered through a client app's own trusted components instead of raw code.

AEO (Answer Engine Optimization)

21 June 2026 · 0:51

Answer Engine Optimization. A name for optimising to appear in AI answer engines.

AGI (Artificial General Intelligence)

26 July 2026 · 1:38

A system that matches or exceeds human performance across the full range of cognitive tasks rather than one narrow slice. There is no agreed definition and no accepted test, so progress is argued through benchmarks that stand in for the term.

AI Agent

6 July 2026 · 1:17

An AI system that plans and executes multi-step tasks autonomously — researching, calling tools, and acting toward a goal rather than just answering.

AI Assistant

6 July 2026 · 1:23

A conversational AI like ChatGPT, Gemini or Claude used for short, varied tasks — usually one question and one answer, not search-style queries.

AI Content Detection

6 July 2026 · 1:17

Classifying whether text was written by AI; as models improve, detection needs constant fine-tuning and hybrid deep-learning-plus-heuristic methods.

AI Crawlers

6 July 2026 · 1:14

The bots AI systems use — training-data scrapers versus agentic assistants — that site owners must decide whether to allow or block.

AI Influence

9 August 2026 · 1:25

The measurable effect content has on what AI systems say: the causal layer beneath AI Visibility, spanning what a model was trained to believe and what it reads at answer time.

AI Mode

6 July 2026 · 1:28

Google's AI-powered search experience — essentially Gemini with search, calculator, time, location and Python tools — answering queries conversationally.

AI Overviews

6 July 2026 · 1:13

Google's AI-generated answer summaries atop search results, which leaked signals suggest are built on its Dialogflow agentic framework.

AI Rank

6 July 2026 · 1:11

Our AI visibility and rank-tracking framework that probes language models to map how they perceive and associate brands.

AI SEO

21 June 2026 · 0:53

A common name for AI-centric SEO: the practice of working toward AI Visibility.

AI Search

2 August 2026 · 1:02

Search where an AI system answers a query directly from retrieved sources, instead of returning a list of links.

AI Visibility

21 June 2026 · 1:03

The broad outcome of AI-centric SEO: being seen, cited, and recommended by AI systems when they answer questions.

AIO (AI Optimization)

21 June 2026 · 0:40

AI Optimization. Another name for the practice of improving AI Visibility.

AP2 (Agent Payments Protocol)

6 July 2026 · 1:04

A payment-agnostic open standard that lets autonomous AI agents authorize and complete transactions securely via cryptographically signed "mandates," removing manual checkout.

Activation Steering

6 July 2026 · 1:42

Injecting a computed direction vector into a model's residual stream at inference time to bias its output toward a target behaviour — without any weight updates or retraining.

Agentic Harness

6 July 2026 · 1:30

The orchestration layer around a language model that equips it with tools, memory, and control flow — turning a text predictor into an agent that can plan, act, and loop.

Alignment

2 August 2026 · 1:05

The effort to make a model behave as intended (helpful, honest, harmless) via RLHF and related methods; it sets a model's defaults about refusals, trust, and which sources it treats as credible.

Annotated Page Content (APC)

5 July 2026 · 1:21

The structured, machine-readable representation Chrome builds from a page's rendering tree when a tab is shared with Gemini. It captures every visible element — text, links, images, forms, tables — as a tree of content nodes, each tagged with geometry, styling, interaction data, and a unique node ID

Anthropic

10 August 2026 · 1:05

The AI safety company behind Claude and the Model Context Protocol, founded 2021. Known for Constitutional AI and for interpretability work on features inside models.

Associative Embeddedness

6 July 2026 · 1:20

A measure of how deeply a brand is woven into a model's memory, from how often and how early the model recalls it across many runs.

Attention

6 July 2026 · 1:18

The transformer mechanism that lets a model weigh how much every token relates to every other, in parallel — the core of modern LLMs.

Binary Classification

6 July 2026 · 1:37

A model task with exactly two possible output labels — yes/no, spam/not-spam, AI/human — the simplest and most common form of supervised classification.

Binary Vector Embeddings

6 July 2026 · 1:35

Compressing float embeddings to one bit per dimension, reducing storage by ~32× with surprisingly small quality loss — useful when speed and scale matter more than marginal precision.

BlockRank

6 July 2026 · 1:33

A method that makes LLM in-context ranking scale to large candidate sets by exploiting the block-sparse structure of attention.

Brand Association Network

6 July 2026 · 1:39

The graph of entities, concepts, sentiments, and competitor brands a language model associates with a given brand, surfaced through systematic probing and tracked over time.

Brand Authority

6 July 2026 · 1:13

How much cognitive space a brand occupies in a model's memory — how readily and prominently it recalls the brand unprompted.

CAPS

6 July 2026 · 1:18

Our proposed Content Attribution Payment Scheme — a framework to fairly attribute and pay creators when AI systems use their content.

Canonical Tag

2 August 2026 · 1:03

A rel=canonical link element naming the preferred URL among duplicates, so ranking signals consolidate onto one address instead of splitting across copies.

Chain of Thought

2 August 2026 · 1:12

A model working through a problem in explicit intermediate steps before answering; reasoning models build this deliberation in rather than relying on the prompt to elicit it.

ChatGPT

10 August 2026 · 1:00

OpenAI's assistant product, launched 30 November 2022. It put a language model in front of a general audience and made conversational answering a surface that competes with search.

Checkpoint

6 July 2026 · 1:25

A saved snapshot of a model's weights at a point during training, enabling recovery, comparison across training stages, and selection of the best-performing version.

Chrome History Embeddings

6 July 2026 · 1:08

Chrome's local system that turns pages you visit into vectors, enabling natural-language search of your own browsing history on-device.

Citation

21 July 2026 · 1:19

A grounding source the model surfaces to the reader as a reference: a subset of what it retrieved, and an attribution of a source rather than a mention of your brand.

Citation Mining

6 July 2026 · 1:09

Systematically running many prompts across AI platforms to record which domains and URLs get cited — mapping the AI citation landscape for a topic.

Class Imbalance

6 July 2026 · 1:21

When one class in a training dataset vastly outnumbers another, causing a naive model to ignore the minority class — requires deliberate handling through weighted loss, resampling, or specialised loss functions.

Classifier

6 July 2026 · 1:27

A model that assigns one or more labels to an input — the fundamental pattern behind query intent models, spam detectors, sentiment models, and content detection.

Claude

10 August 2026 · 1:06

Anthropic's model family, first released March 2023, trained with Constitutional AI and named in Haiku, Sonnet and Opus sizes since the Claude 3 generation.

Clustering

10 August 2026 · 1:39

Grouping items by similarity with no labels. Run over embeddings it produces topic groups, near-duplicate sets and query groups that nobody defined in advance.

Content Fetcher

6 July 2026 · 1:03

One of AI Mode's two page-reading tools — it retrieves a batch of structured sources at once, versus browsing a single URL.

Content Freshness

2 August 2026 · 1:04

How recent and actively updated a page's content is; it weighs on time-sensitive queries and decides which sources an AI retrieves when it runs a live search past its knowledge cutoff.

Content Node

5 July 2026 · 1:17

A single element in Chrome's page representation. Each node carries a type (one of 21, from headings and links to form controls and dialogs), bounding box coordinates, text styling, accessibility metadata, and an ID that the AI uses to reference or interact with that specific element.

Content Substance Classification

6 July 2026 · 1:23

Detecting whether content actually says something or just sounds like it does — separating genuine substance from eloquent fluff.

Context Engineering

6 July 2026 · 1:20

Deliberately constructing the semantic environment of a prompt to activate specific representational circuits within a model — moving beyond keyword targeting toward architecture-aware content design.

Context Window

6 July 2026 · 1:26

The maximum number of tokens a model can read at once — the total working memory available for a prompt, grounding documents, conversation history, and the generated response combined.

Cosine Similarity

6 July 2026 · 1:03

A measure of how alike two vectors are by the angle between them; for normalised vectors it's identical to a dot product.

Deep Learning

10 August 2026 · 1:17

Machine learning with many stacked layers of learned transformations, trained end to end by gradient descent, so features are learned from data rather than designed by hand.

Digital Marketing

10 August 2026 · 1:19

The set of channels that reach an audience through digital media. Answer engines are the newest one and break the measurement model the others share, because there is often no click.

Document Grounding

6 July 2026 · 1:04

Grounding an AI answer in the content of a specific web page you supply, rather than in results from an open web search.

DocumentChunker

6 July 2026 · 0:58

Chrome's internal Blink algorithm that recursively walks a page's DOM layout tree to group text nodes into semantically cohesive passages of up to 200 words.

DomDistiller

6 July 2026 · 1:15

Chrome's Reader Mode engine, whose content-extraction heuristics are a public proxy for how machines separate main content from boilerplate.

Dynamic Retrieval

6 July 2026 · 1:08

Google's mechanism for deciding, per query, whether Gemini runs a live search — governed by a default grounding threshold of 0.3.

E-E-A-T

2 August 2026 · 1:13

Experience, Expertise, Authoritativeness, and Trustworthiness: the qualities Google's rater guidelines use to judge whether content and its author can be relied on.

Early Stopping

6 July 2026 · 1:30

Halting training when validation performance stops improving rather than running for a fixed number of epochs — the primary defence against overfitting in practice.

EmbeddingGemma

6 July 2026 · 1:10

Google's compact open embedding model, a miniaturised relative of Gemini, that turns text into meaning-carrying vectors for search understanding.

Entity

2 August 2026 · 0:57

For our AI SEO work, the smallest irreducible thing a client wants to be visible for, treated as a cluster of the many query and prompt variants that stem from it.

Epoch

6 July 2026 · 1:29

One complete pass through the entire training dataset — the basic unit of training progress, used to track how many times the model has seen each example.

F1 Score

6 July 2026 · 1:31

The harmonic mean of precision and recall — the standard single-number summary of classifier performance when both false positives and false negatives matter.

Fine-tuning

6 July 2026 · 1:33

Continuing training of a pre-trained model on a smaller, task-specific dataset to specialise its behaviour — the standard way DEJAN adapts base models into production classifiers.

Focal Loss

6 July 2026 · 1:28

A modified cross-entropy loss that down-weights easy, well-classified examples so training focuses on hard cases — the standard solution when positive examples are rare.

Frequency

21 June 2026 · 0:41

How often your citations and mentions occur across AI queries and over time.

GEO (Generative Engine Optimization)

21 June 2026 · 1:07

Generative Engine Optimization. A name for optimising to appear in generative AI engines.

Gemini

10 August 2026 · 1:15

Google DeepMind's natively multimodal model family, announced December 2023, and the model line behind AI Overviews and AI Mode in Google Search.

Generate-then-Ground

6 July 2026 · 1:16

Google's approach of drafting an answer first, then searching for sources to attribute it — which we argue is more fragile than retrieving first.

Generative Engine

2 August 2026 · 1:07

An umbrella term for AI answer systems, coined to give Generative Engine Optimization a target; it conflates search engines, assistants, and agents and names no actual product.

Generative self-retrieval

27 June 2026

A reasoning model surfacing facts from its own parameters by generating them as text, where those generated facts then condition and improve its final answer.

Google Search Console

10 August 2026 · 1:19

Google's free service for a verified site owner: the queries it appeared for with impressions, clicks and average position, plus indexing, crawl and Core Web Vitals status.

Grounding

21 June 2026 · 0:40

The process by which an AI system answers with current information: it runs a search, retrieves pages, reads short extracts, and writes its answer from them rather than from memory.

Grounding Bias

21 June 2026 · 1:08

A form of secondary bias: how much an AI model defers to, or discounts, retrieved sources once they are grounded into the context, rather than leaning on what it already believed.

Grounding Chunk

6 July 2026 · 1:19

The slice of a page's content an AI system actually uses to support a sentence, drawn from a fixed per-query word budget.

Grounding Classifier

6 July 2026 · 1:13

Our model that predicts whether a query "deserves grounding" — whether an AI system will run a live web search to answer it.

Grounding Snippet

21 June 2026 · 0:55

A short, extractive selection of sentences pulled from a page and supplied to an AI model as the evidence it grounds an answer on; the atomic unit of visibility in AI search.

Hallucination

2 August 2026 · 1:16

When a large language model produces fluent, confident text that is false or unsupported, because it predicts probable wording rather than drawing on verified facts.

Hyperparameters

6 July 2026 · 1:21

Settings configured before training begins that control how a model learns — learning rate, batch size, epoch count, dropout rate — as distinct from the weights the model learns itself.

Implicit Queries

6 July 2026 · 1:13

Searches an AI system formulates and runs on your behalf without you typing them, based on your on-device context and profile.

In-Context Ranking

6 July 2026 · 1:02

Re-ranking candidate documents for a query by feeding them all into an LLM's context and letting it pick the most relevant.

Inference

6 July 2026 · 1:40

Running a trained model on new inputs to produce predictions — the production mode of a model, as distinct from the training phase where weights are learned.

Information Density

2 August 2026 · 1:07

How much genuine information content carries per unit of length: the ratio of substance to words, the axis our Cyberfluff substance work scores.

Information Gain

2 August 2026 · 1:08

How much new information a document adds beyond what is already known; defined in information theory as a drop in entropy, and applied in search to reward content that adds rather than repeats.

Information Retrieval

10 August 2026 · 1:31

The field behind search: matching a query to documents in a collection and ranking them by estimated relevance. It supplies the retrieval half of RAG and of grounding in AI answers.

Keyword Research

10 August 2026 · 1:03

Finding the queries an audience uses, sizing the demand and mapping them to pages. In AI search the unit shifts from a keyword to a prompt and the queries it fans out into.

Knowledge Cutoff

2 August 2026 · 0:53

The date after which a model has no built-in knowledge, because its training data stops there; recent facts are unknown unless supplied through grounding or retrieval.

Knowledge Graph

2 August 2026 · 0:57

A structured network of entities and their relationships stored as nodes and edges; the explicit kind powers Google's knowledge panels, and an LLM carries an implicit statistical version in its weights.

Kolmogorov Complexity

26 July 2026 · 1:34

The length of the shortest program that reproduces a given string, a measure of how much information one specific object carries, approximated in practice by how well it compresses.

LLM Search Volume Prediction

6 July 2026 · 1:20

Using a language model to estimate a query's monthly search volume — directionally useful for sizing topics, but not precise.

Language Model

6 July 2026 · 1:19

A model trained to understand and generate text by learning the statistical patterns of language — the foundational class of model underpinning search, AI assistants, and every DEJAN classifier.

Large Language Model

6 July 2026 · 1:23

A language model at sufficient scale — billions of parameters, trained on web-scale text — to exhibit emergent capabilities like reasoning, instruction following, and open-ended generation. Gemini, GPT, Claude, and Llama are all LLMs.

Latent Entities

6 July 2026 · 1:18

A critical piece of information a title deliberately withholds — the subject, reason, process or outcome — to force a click.

LoRA

6 July 2026 · 1:39

Low-Rank Adaptation — a parameter-efficient fine-tuning technique that trains small adapter matrices instead of updating a model's full weights, dramatically reducing memory and compute requirements.

LongFormContextResult

5 July 2026 · 1:18

The data wrapper Chrome uses to deliver extracted page content to Gemini's context window. It separates external source material from conversation, ensuring the model grounds its answers in the page rather than its training data. Each block carries a session-wide index used for citation mapping.

Loss Function

6 July 2026 · 1:34

The mathematical measure of how wrong a model's predictions are — the quantity training minimises, and the choice of which shapes everything about how the model learns.

MUVERA

6 July 2026 · 1:24

A Google method for multi-vector retrieval that compresses many per-token vectors into a single fixed-dimensional encoding for fast search.

Machine Learning

10 August 2026 · 1:25

Fitting a model's parameters to data so it performs a task no one wrote rules for. The training run adjusts parameters against a loss; the fitted parameters carry the behaviour into inference.

Manifold Hypothesis

6 July 2026 · 1:03

The principle that natural data lies on a low-dimensional manifold within high-dimensional space — the theoretical foundation for why vector embeddings capture meaning rather than noise.

Matryoshka Representation Learning

6 July 2026 · 1:13

A training technique that nests coarse-to-fine information inside one embedding, so it can be truncated to fewer dimensions with little loss.

Mechanistic Interpretability

6 July 2026 · 1:16

Techniques for tracing what happens inside a model — logging activations and patching them — to see how and where it forms an output.

Mention

21 July 2026 · 1:00

Your brand appearing by name in the readable text of an AI answer: a generation outcome that leans on the model's prior more than on which pages were grounded, and the one that carries commercial weight.

Mixture of Experts

6 July 2026 · 1:41

An architecture where only a subset of specialised sub-networks (experts) activate for each token, allowing a model to have enormous total capacity while spending only a fraction of that compute on any single input.

Model Context Protocol

6 July 2026 · 1:23

An open protocol that gives AI models a standard way to connect to external tools and data sources; WebMCP brings the idea to websites.

Model Distillation

6 July 2026 · 1:22

Training a smaller student model to replicate the behaviour of a larger teacher model, producing compact models that punch above their parameter count.

Model Probing

6 July 2026 · 1:18

Systematically querying a language model with structured prompts to map its latent biases, brand associations, and internal knowledge graphs without relying on traditional rank trackers.

Multi-label Classification

6 July 2026 · 1:42

A classification task where each input can be assigned more than one label simultaneously — a query can be both informational and local; a sentence can be both positive and sarcastic.

Multimodal Model

6 July 2026 · 1:30

A model that processes and reasons across more than one type of input — text, images, audio, video — within a single unified architecture rather than routing each modality to a separate specialist.

Named Entity Recognition

2 August 2026 · 1:08

The NLP task of finding named things in text and classifying them as people, organisations, locations, dates, and products; the traditional pipeline sense of an entity.

Natural Language Processing (NLP)

10 August 2026 · 1:23

The field concerned with computers reading and producing human language. Its pipeline of tokenising, tagging and parsing has largely been absorbed into general-purpose language models.

Needle in a Haystack

7 July 2026 · 1:31

A stress test that plants a single fact deep inside a long context and checks whether the model can retrieve it — the standard way to measure whether a large context window is genuinely usable, or just advertised.

Neural Network

10 August 2026 · 1:09

Layers of weighted sums passed through a non-linear function, with the weights fitted by gradient descent. The base structure under every current AI model.

Node ID

5 July 2026 · 1:33

A sequential identifier assigned to each content node during extraction via depth-first traversal. These IDs appear in the structured Markdown output as {#ID} references, enabling the AI to target specific elements on the page for reading, citing, or interaction.

Number of Citations

21 June 2026 · 0:38

The absolute count of times your pages are cited in AI answers.

Number of Mentions

21 June 2026 · 0:31

The absolute count of times your brand is named in AI answers.

On-Device Models

6 July 2026 · 1:13

Machine-learning models that run locally in software like Chrome — fast, private, and shaping how your content is read before it reaches a server.

Open Knowledge Format

21 June 2026 · 0:38

An open, vendor-neutral format for packaging the knowledge AI systems need, as plain markdown files any model or agent can read.

OpenAI

10 August 2026 · 1:04

The company behind the GPT model family and ChatGPT. Founded 2015; its GPT-3 and InstructGPT results established in-context learning and RLHF as standard practice.

Otsu's Thresholding (for Queries)

6 July 2026 · 1:06

An unsupervised method borrowed from computer vision that sets optimal per-label confidence cut-offs in query classification by maximizing the variance between positive and negative score clusters.

Overfitting

6 July 2026 · 1:29

When a model memorises training examples rather than learning generalisable patterns — it scores well on training data but poorly on new inputs it hasn't seen.

Parametric Memory

10 August 2026 · 1:29

Knowledge held in a model's own weights and recalled at inference with no lookup, as distinct from non-parametric memory, which is text retrieved into the context window at request time.

Perplexity

2 August 2026 · 1:10

How well a model predicts text: the exponential of average per-token cross-entropy, tied to entropy, and the standard intrinsic metric for evaluating language models. Not the search company of the same name.

Pre-training

6 July 2026 · 1:31

Training a model from scratch on a large general corpus to build broad language understanding before any task-specific work begins.

Precision

6 July 2026 · 1:25

Of all the times a model predicted positive, the fraction it was actually right — a measure of how trustworthy positive predictions are.

Primary Bias

21 June 2026 · 1:04

An AI model's ungrounded confidence in an entity, formed during training and present before any retrieval; in AI search it is the largest single factor in whether a source is selected.

Prompt

2 August 2026 · 0:54

The text given to a large language model to condition its response: the system instruction, context, retrieved documents, and user message the model reads before it generates.

Prompt Engineering

10 August 2026 · 1:32

Writing and structuring model input so the output is reliable enough to use. Where the problem is what the model is given rather than how it is asked, the discipline is context engineering.

Prompt Injection

2 August 2026 · 1:20

An attack that smuggles instructions into the text a model reads so it obeys them; the indirect form hides them in retrieved web pages or documents an AI agent acts on.

Quadratic Attention Cost

7 July 2026 · 1:27

The property that standard transformer attention grows with the square of the input length — doubling the context roughly quadruples the compute and memory — which is why long context windows are expensive and why grounding budgets exist.

Quantization

10 August 2026 · 1:28

Storing weights and activations at lower numeric precision. Going from 16-bit to 4-bit cuts memory roughly fourfold and raises throughput, at a measurable accuracy cost.

Query Classification

6 July 2026 · 1:19

Assigning search queries to labels — like intent categories — so they can be routed, mapped to content, and reported at scale.

Query Deserves Grounding

6 July 2026 · 1:07

Google's internal decision about whether a query needs a live web search; below its threshold, the model answers from memory.

Query Fanout

21 June 2026 · 0:44

The step where an AI system splits one prompt into several single-intent sub-queries, each retrieving its own sources; a page can be grounded for one angle of a question and absent for another.

Query Well-formedness

6 July 2026 · 1:20

The property of a search query being grammatically natural and unambiguous — a signal used to distinguish full-sentence questions from fragmentary keyword strings, with implications for query expansion and intent classification.

RLHF

2 August 2026 · 1:15

The training stage that aligns a model to human preferences using a reward model trained on human rankings; the main lever behind alignment, distinct from fine-tuning on labelled examples.

Rank

21 June 2026 · 0:47

How prominently your citations and mentions appear within an AI answer.

Reasoning Model

2 August 2026 · 1:04

A language model trained to work through an extended internal chain of thought and spend more compute on harder problems before answering; stronger on multi-step tasks.

Recall

6 July 2026 · 1:12

Of all the actual positive cases in the data, the fraction the model correctly identified — a measure of how much the model misses.

Recursive Self-Improvement

2 August 2026 · 1:22

An AI improving its own capabilities and using each improved version to improve itself again, in a compounding loop; the mechanism behind the intelligence-explosion argument toward AGI.

Relevance Engineering

21 June 2026 · 0:58

Engineering a page's semantic relevance to a query with embeddings and vector math, treating visibility as an engineering problem rather than keyword optimization.

Retrieval-Augmented Generation

6 July 2026 · 1:13

A framework that retrieves relevant passages first, then feeds query plus evidence to the model so it answers from real text.

Reverse Prompting

6 July 2026 · 1:28

Reconstructing the most likely prompt that generated a piece of AI output — used for detection, content auditing, and understanding how models respond to specific framings.

Rufus

6 July 2026 · 1:19

Amazon's AI shopping assistant — a multi-component RAG system that plans a query, retrieves across Amazon's own sources, and streams an answer.

Search Engine Optimization (SEO)

2 August 2026 · 1:16

The practice of improving a page's visibility in unpaid search results through technical, content, and authority work; its foundations still underpin AI visibility because AI systems retrieve from the same web.

Search Intent

2 August 2026 · 1:08

The underlying goal behind a query or prompt (informational, navigational, transactional, or commercial), which a system optimises to satisfy rather than the exact words used.

Secondary Bias

21 June 2026 · 0:58

The post-retrieval layer of selection: how an AI model treats, weights, and is swayed by content once it has been retrieved. Unlike primary bias, it is addressable now.

Selection Rate

21 June 2026 · 0:50

The rate at which an AI system chooses a given source from the available grounding candidates when composing an answer; the AI-search equivalent of click-through rate.

Selection Rate Optimization

21 June 2026 · 0:44

The AI-search counterpart to CTR optimization: improving how often AI systems select your content as the grounding source for their generated answers.

Semantic Compression

6 July 2026 · 0:58

Writing dense, self-contained passages so that when an AI system extracts an isolated grounding chunk, the core context like the brand or product name travels with it.

Semantic Search

2 August 2026 · 1:13

Search that matches a query to content by meaning rather than exact keywords, comparing vector embeddings; the retrieval method behind AI answers.

Shannon Entropy

26 July 2026 · 1:37

The average number of bits a source needs per symbol, computed from the probability distribution over its symbols. It sets the floor no lossless compressor can beat, and it is the unit language model quality is reported in.

Share of Citations

21 June 2026 · 0:50

The proportion of cited sources in AI answers that are yours.

Share of Mentions

21 June 2026 · 1:10

The proportion of brand mentions in AI answers that are yours.

Share of Voice

21 June 2026 · 0:42

Your overall slice of the AI answer space for a topic, relative to competitors.

Site Engagement Score

6 July 2026 · 1:28

Chrome's internal 0–100 per-domain user engagement metric, built from navigation events, input interactions, and media consumption, decaying over time when a site is not visited.

Site Search in AI Mode

6 July 2026 · 1:10

Vertex AI's website-search capability that lets an owner run an AI-Mode-style search over their own content, confirming the Prepare→Retrieve→Signal→Serve flow.

Source Authority

2 August 2026 · 1:13

How trustworthy and credible a source is judged to be, which shapes whether search engines and AI systems retrieve, rely on, and cite it.

Structured Data (Schema Markup)

2 August 2026 · 1:03

Machine-readable annotation (schema.org, usually in JSON-LD) that states what a page's content is, so engines read defined entities and facts instead of inferring them from prose.

Synthetic Training Data

6 July 2026 · 1:20

AI-generated datasets used to train or fine-tune language models; in AI SEO, a potential early-stage strategy for influencing which brands and associations future models learn.

System Prompt

2 August 2026 · 1:01

The standing instruction at the top of a model's context that sets its role, rules, and behaviour for a whole conversation, given priority over user messages and usually hidden.

Technical SEO

10 August 2026 · 1:28

Whether a site can be crawled, rendered, indexed and served fast enough to compete. A precondition rather than a tactic: a page outside the index cannot rank or be cited.

Temperature

6 July 2026 · 1:26

A setting that rescales next-word probabilities to control randomness — low values make output focused, high values more diverse.

Token

6 July 2026 · 1:30

The atomic unit a language model reads and generates — a subword chunk, not a whole word — that determines how text is measured, priced, and processed.

Token Probability

6 July 2026 · 1:07

The likelihood a model assigns to each candidate next token; low probabilities flag where the model is uncertain about your brand or topic.

Tokenization

6 July 2026 · 1:11

Breaking text into subword tokens a model can process; each token carries a frequency bias from how often it appeared in training.

Tokenizer

6 July 2026 · 1:39

The component that converts raw text into the sequence of tokens a model actually processes — fixed at pre-training time, and different for every model family.

Tool Use

2 August 2026 · 1:02

How a model acts beyond text: it emits a structured call to an external tool such as search, code, or an API; a harness runs it and feeds the result back. The mechanism under AI agents.

Top-p Sampling

6 July 2026 · 1:24

A decoding method that samples the next word from the smallest set of top candidates whose probabilities sum to a threshold p.

Training Data

6 July 2026 · 1:39

The labelled or unlabelled examples a model learns from — the single biggest determinant of what a model knows, what it believes, and which brands and entities it associates with which topics.

Transfer Learning

6 July 2026 · 1:33

Applying knowledge learned on one task or dataset to a different but related task — the principle that makes fine-tuning a pre-trained model far more efficient than training from scratch.

Transformer

10 August 2026 · 1:34

The architecture behind current language models: stacked self-attention and feed-forward blocks that process a whole sequence in parallel, introduced by Vaswani et al. in 2017.

Tree Walker

6 July 2026 · 1:20

Our brand-analysis tool that reveals how Gemini talks about a brand by walking the probabilistic paths of the model's own language.

URL Context Tool

6 July 2026 · 1:14

Google's Gemini capability that fetches and reads the content of a specific supplied URL, enabling document grounding.

Vector Embedding

6 July 2026 · 1:02

A dense numeric representation of text that captures meaning and intent, so machines can compare content by similarity rather than keywords.

Vector Embedding Optimization

6 July 2026 · 1:18

Choosing embedding dimension and quantisation to balance speed, storage and accuracy — often gaining efficiency with little loss.

Vector Search

6 July 2026 · 1:11

Finding the most relevant content by comparing meaning-carrying vectors, ranking by nearest neighbour rather than keyword overlap.

Vertex AI

10 August 2026 · 1:14

Google Cloud's managed platform for running, tuning and serving models, and the enterprise route to the Gemini API, with grounding, batch prediction, tuning and a model garden.

Web Search Grounding

6 July 2026 · 1:18

The process where an AI model runs a live web search, reads the results, and weaves some of those pages into its answer.

WebMCP

6 July 2026 · 1:13

A proposed web standard that lets sites expose structured tools to AI agents, so agents invoke defined actions instead of screen-scraping the page.

llms.txt

2 August 2026 · 0:52

A proposed root file offering a markdown map of a site's key pages for language models; put forward in 2024 but not adopted by Google or other major AI platforms.