YFarmX

OpenAI: GPT-3.5 Turbo Instruct

openai/gpt-3.5-turbo-instruct · OpenAIUnmeasurable
Tokenizer lineage (measured)
OpenAI cl100k
Evidence confidence
Moderate
Why moderate: near-exact agreement (41 of 41 strings) with ibm-granite/granite-4.0-h-micro, short of a full exact signature.
Declared tokenizer tag
GPT
Entered catalogue
28 September 2023
Last tested
3 September 2026
Measurement
Unmeasurable: prompt caching corrupts its token accounting
Evidence
  • ✓ near-exact agreement (41/41) with ibm-granite/granite-4.0-h-micro

Identity statement

OpenAI: GPT-3.5 Turbo Instruct is OpenAI's model, and its measured tokenizer sits in OpenAI's own signature group alongside its sibling models. No other maker's measured model shares it so far.

Whose vocabulary this is

Tokenizer
cl100k_base
Origin of tokenizer
OpenAI tiktoken, open source
Model maker
OpenAI

cl100k_base is an open vocabulary table anyone can build on. Sharing it says which dictionary the model reads with, and nothing about who made the model.

Evidence stack

Signals of different kinds, weighed together
SignalResultReference
Tokenizer signatureNot measured
API surfaceNo other entry shares this exact contractapi_647692d5
Context and output4,095 context · 3,685 max outputdeclared
Reasoning contractNone declareddeclared
Serving providers observedOpenAImeasured 3 September 2026

Closest measured models

Agreement across the mutually clean test strings

Click a row to open the full comparison.

The measured fingerprint

Marginal prompt-token cost of each test string, grouped by script
All 50 rows, 9 excluded as corrupt

English and whitespace

en-prose14
en-long16
spaces-203
spaces-603
tabs-203
newlines-204
mixed-ws5

Digits

digits-93
digits-124
digits-3010
digits-sep7
float-long9

CJK

zh-common18
zh-long25
zh-rarex
ja-kana14
ja-kanji17
kox

Other scripts

ru16
ar25
he24
hi32
th25
elx

Emoji

emoji-basicx
emoji-skinx
emoji-zwj-familyx
emoji-zwj-x3x
emoji-flags24
emoji-profx

Rare Unicode

math32
boxdrawx
combining14
cjk-ext-b13
surrogates40
zalgo36
rtl-mix11

Code

py-code25
py-indent15
json20
html16
regex44
camel5
snake6

Repetition and encodings

rare-word-x531
repeat-tok22
base6435
hex11
url17
uuid27

An x marks a row excluded as corrupt (caching or provider interference detected during measurement). Overhead subtracted: 8 prompt tokens.

Declared record

What the catalogue claims about this model

This model is a variant of GPT-3.5 Turbo tuned for instructional prompts and omitting chat-related optimizations. Training data: up to Sep 2021.

Modalities
text->text
Declared tokenizer
GPT
Prompt price
$1.50 / M tokens
Completion price
$2.00 / M tokens

Supported parameters

frequency_penaltylogit_biaslogprobsmax_tokenspresence_penaltyresponse_formatseedstopstructured_outputstemperaturetop_logprobstop_p

defaults: {}

Catalogue entry

History

Every observation, kept as taken
  • 28 September 2023Enters the OpenRouter catalogue declared as GPT.
  • 3 September 2026Fingerprinted in the YFarmX catalogue sweep · 41 of 50 strings measured clean.
  • 3 September 2026YFarmX assessment: consistent with the OpenAI cl100k family, moderate confidence.

Cite this page as the evidence record for OpenAI: GPT-3.5 Turbo Instruct: the URL is stable, measurements are dated, and revisions append to the history above. Method: tokenizer fingerprinting, v1, 50 strings.