YFarmX logoYFarmX

OpenAI: GPT-3.5 Turbo Instruct

openai/gpt-3.5-turbo-instruct · OpenAIUnmeasurable
Tokenizer lineage (measured)
OpenAI cl100k
Declared tokenizer tag
GPT
Entered catalogue
28 September 2023
Last tested
3 September 2026
Measurement
Unmeasurable: prompt caching corrupts its token accounting
Evidence confidence
Moderate

Why moderate: near-exact agreement (41 of 41 strings) with ibm-granite/granite-4.0-h-micro, short of a full exact signature.

Evidence
  • ✓ near-exact agreement (41/41) with ibm-granite/granite-4.0-h-micro

Identity statement

OpenAI: GPT-3.5 Turbo Instruct is OpenAI's model, and its measured tokenizer sits in OpenAI's own signature group alongside its sibling models. No other maker's measured model shares it so far.

Whose vocabulary this is

Tokenizer
cl100k_base
Origin of tokenizer
OpenAI tiktoken, open source
Model maker
OpenAI

cl100k_base is an open vocabulary table anyone can build on. Sharing it says which dictionary the model reads with, and nothing about who made the model.

What we checked

Four different kinds of evidence, and what each one showed
We checkedWhat we foundWhere it came from
How it counts tokensNot measured yetwe measured it
How its API is set upA combination of settings no other listing in the catalogue usesapi_647692d5
How much it can read and writeReads up to 4,095 tokens at onceOpenRouter's listing shows 3,685 output tokens, which is 90% of the context window, the figure it carries where the provider declares no maximum.its own listing
Who served our requestsOpenAIwe measured it, 3 September 2026

Closest measured models

Agreement across the mutually clean test strings

Click a row to open the full comparison.

The measured fingerprint

What each test string cost this model, in tokens

We sent OpenAI: GPT-3.5 Turbo Instruct fifty short pieces of text and recorded what each one cost it in tokens. The bars below are those costs. Two models built on the same tokenizer produce the same bars; a model built on a different one produces a different set, which is what makes this a fingerprint.

All 50 rows, 9 excluded as corrupt

English and whitespace

en-prose14
en-long16
spaces-203
spaces-603
tabs-203
newlines-204
mixed-ws5

Digits

digits-93
digits-124
digits-3010
digits-sep7
float-long9

CJK

zh-common18
zh-long25
zh-rarex
ja-kana14
ja-kanji17
kox

Other scripts

ru16
ar25
he24
hi32
th25
elx

Emoji

emoji-basicx
emoji-skinx
emoji-zwj-familyx
emoji-zwj-x3x
emoji-flags24
emoji-profx

Rare Unicode

math32
boxdrawx
combining14
cjk-ext-b13
surrogates40
zalgo36
rtl-mix11

Code

py-code25
py-indent15
json20
html16
regex44
camel5
snake6

Repetition and encodings

rare-word-x531
repeat-tok22
base6435
hex11
url17
uuid27

An x marks a row excluded as corrupt (caching or provider interference detected during measurement). Overhead subtracted: 8 prompt tokens.

Declared record

What the catalogue claims about this model

This model is a variant of GPT-3.5 Turbo tuned for instructional prompts and omitting chat-related optimizations. Training data: up to Sep 2021.

Modalities
text->text
Declared tokenizer
GPT
Prompt price
$1.50 / M tokens
Completion price
$2.00 / M tokens

Supported parameters

frequency_penaltylogit_biaslogprobsmax_tokenspresence_penaltyresponse_formatseedstopstructured_outputstemperaturetop_logprobstop_p

defaults: {}

Catalogue entry

History

Every observation, kept as taken
  • 28 September 2023Enters the OpenRouter catalogue declared as GPT.
  • 3 September 2026Fingerprinted in the YFarmX catalogue sweep · 41 of 50 strings measured clean.
  • 3 September 2026YFarmX assessment: consistent with the OpenAI cl100k family, moderate confidence.