Site icon Ampersand Tutorials

Python Object Model: Names, References and Mutation

Quick answer: Python variables are names bound to objects, not boxes holding values. Assignment never copies; it points a second name at the same object. is compares identity (same object in memory) while == compares value, small integers from -5 to 256 are cached and shared, and list += mutates in place where tuple += rebinds a new object. Every surprising Python bug — mutable defaults, aliased lists, “spooky” shared state — follows from these four rules.

Part 10 of our Python series — the first article of Module 3, where we go past syntax into how CPython actually manages your data. Start at Module 1 if you are new, or read Module 2 for the standard library.

Names are labels, not containers

In C an int x reserves a box; in Python x = [1, 2, 3] creates a list object somewhere in memory and attaches the name x to it. The name is a key in a namespace dictionary, the object is separate.

a = [1, 2, 3]
b = a                 # no copy: two names, one list
print(a is b, id(a) == id(b))   # True True
b.append(4)
print(a)              # [1, 2, 3, 4] -- a "changed" because a IS b

id() returns the object’s identity — its memory address for the CPython implementation. is is exactly the id() comparison, fast and reliable. This is why senior developers say Python has no variables in the C sense: it has names bound to objects, and the binding can be redirected at any time.

The moment you write b = a[:], list(a), copy.copy(a), or a + [], you create a new object and the two names diverge. Every Python bug of the form “I copied the list but the original still changed” is a case where a copy was shallow and the inner objects were still shared.

is vs ==: identity versus equality

== invokes __eq__ and asks “are these equal?”; is asks “is this literally the same object?” They coincide for singletons and cached integers and diverge everywhere else.

[1, 2] == [1, 2]     # True  -- equal contents
[1, 2] is [1, 2]     # False -- two distinct list objects

x = int("256"); y = int("256")
x is y               # True  -- CPython caches -5..256
x = int("257"); y = int("257")
x is y               # False -- outside the cache, two objects

The small-integer cache exists because programs create small ints constantly; pre-allocating 262 of them is cheaper than allocating repeatedly. Integers outside that range, freshly built strings, and most objects are distinct even when equal. Constants written literally in the same code object are folded into one object — compile 257 is 257 and inspect the code object:

def f():
    return 257 is 257

f.__code__.co_consts    # (257,) -- one shared constant

That is a compiler optimisation, not a language guarantee, which is exactly why relying on is for numbers is wrong. Python 3.14 now even warns you: writing x is 257 raises a SyntaxWarning telling you to use ==.

Use is only for singletons: None, True, False, NotImplemented, Ellipsis. if x is None is correct and idiomatic; if x == None is a code smell that breaks for objects with unusual __eq__.

Strings have their own sharing mechanism — interning. Identifier-like literals are interned automatically, but runtime-built strings are not:

import sys
s1 = "".join(["py", "thon"]); s2 = "".join(["py", "thon"])
s1 is s2                              # False
sys.intern(s1) is sys.intern(s2)      # True -- explicit interning

Interning pays off when you hold millions of repeatedly-seen strings: identical keys collapse to one object, so dict lookups compare a pointer first and skip character-by-character comparison.

Mutation versus rebinding: the += distinction

list += other calls list.__iadd__, mutating in place. tuple += other has no __iadd__, so Python falls back to __add__ and rebinds the name to a brand-new tuple. Same operator, entirely different consequences for anyone else holding a reference.

t = (1, 2); alias = t
t += (3,)
print(t, alias)        # (1, 2, 3) (1, 2)      -- alias untouched

l = [1, 2]; alias = l
l += [3]
print(l, alias)        # [1, 2, 3] [1, 2, 3]   -- alias sees the change

This is the root of the mutable default argument trap, the single most common interview question in Python:

def bad(item, acc=[]):      # the list is created ONCE, at def time
    acc.append(item)
    return acc

bad(1), bad(2), bad(3)      # ([1, 2, 3], [1, 2, 3], [1, 2, 3])

The default object lives on bad.__defaults__ and every call without acc= appends to the same list. The fix is the None sentinel:

def good(item, acc=None):
    if acc is None:
        acc = []
    acc.append(item)
    return acc

good(1), good(2), good(3)   # ([1], [2], [3])

Top-1% habit: default arguments are evaluated once and shared, so never default to a mutable object — and the same rule explains why def f(now=datetime.now()) freezes the timestamp at import time.

Shallow copy versus deep copy

OperationCopiesNested objects shared?
b = anothingyes (it is the same object)
a[:], list(a), copy.copy(a)outer containeryes
copy.deepcopy(a)container and everything insideno
import copy
inner = [1, 2]
shallow = copy.copy([inner])
deep = copy.deepcopy([inner])

inner.append(3)
print(shallow)   # [[1, 2, 3]]  -- leaked
print(deep)      # [[1, 2]]     -- isolated

deepcopy is correct but slow: it walks the whole graph, and memoises to survive cycles. Reach for it when a structure is genuinely nested or shared; a list of tuples or strings only needs copy.copy, because immutable elements cannot leak changes.

Measuring the real cost of attributes

Every normal instance carries a __dict__ — a hash table per object. For a class with two fields that overhead dominates the data. __slots__ replaces the per-instance dict with fixed slots:

import tracemalloc
import sys

class DictAttrs:
    def __init__(self, x, y): self.x, self.y = x, y

class Slotted:
    __slots__ = ("x", "y")
    def __init__(self, x, y): self.x, self.y = x, y

for cls in (DictAttrs, Slotted):
    tracemalloc.start()
    objs = [cls(i, i) for i in range(100_000)]
    _, peak = tracemalloc.get_traced_memory()
    tracemalloc.stop()
    # note: getsizeof reports the SAME size for both classes -- it cannot see
    # the separate __dict__ object, which is why tracemalloc is used here
    print(f"{cls.__name__:10s} {peak / 1024 / 1024:5.2f} MiB  "
          f"getsizeof={sys.getsizeof(objs[0])}")
    del objs

Measured on CPython 3.14, 100,000 instances of a two-field class:

ClassPeak memoryNotes
normal (__dict__)12.20 MiBattribute names stored per instance
__slots__8.38 MiB31% less memory, faster creation

__slots__ also makes attribute lookup faster and prevents typos creating accidental attributes. The costs: no vars(), no dynamic attributes, and multiple inheritance with non-empty slots from two parents is awkward. Since Python 3.10 the modern spelling is a decorator:

from dataclasses import dataclass

@dataclass(frozen=True, slots=True)
class Point:
    x: float
    y: float

That gives you generated __init__/__repr__/__eq__, immutability, and the slot memory profile in three lines.

The hash and equality contract

Objects used as dict keys or set members must be hashable — hash must stay constant while the object is in the container, and equal objects must have equal hashes. Lists and dicts are unhashable; tuples are hashable only if every element is, which is why (1, (2, 3)) can be a dict key but ([1], 2) cannot.

If you define __eq__ on a class, Python sets __hash__ to None and the class becomes unhashable. For immutable value objects, define both:

class Money:
    __slots__ = ("amount", "currency")
    def __init__(self, amount, currency):
        self.amount, self.currency = amount, currency
    def __eq__(self, other):
        return (self.amount, self.currency) == (other.amount, other.currency)
    def __hash__(self):
        return hash((self.amount, self.currency))

Never mutate a field that participates in __hash__ after putting the object in a set or dict — the object lands in the wrong bucket and effectively disappears.

Complete executable example

This script is a self-checking tour of every rule above: it asserts identities, proves the copy semantics, and reports real memory numbers on your own machine.

# object_model.py -- identity, aliasing, copying, and attribute cost
import copy
import sys
import tracemalloc


def rule(title):
    print("\n" + title)
    print("-" * len(title))


rule("1. assignment binds a second name, it does not copy")
a = [1, 2, 3]
b = a
b.append(4)
print("a is b:", a is b, "| a =", a)
assert a is b and a == [1, 2, 3, 4]

rule("2. identity versus equality")
print("[1,2] == [1,2]:", [1, 2] == [1, 2])
print("[1,2] is [1,2]:", [1, 2] is [1, 2])
print("int('256') is int('256'):", int("256") is int("256"))
print("int('257') is int('257'):", int("257") is int("257"))

rule("3. += mutates a list but rebinds a tuple")
t = (1, 2)
t_alias = t
t += (3,)
print("tuple:", t, t_alias, "| same object:", t is t_alias)
lst = [1, 2]
l_alias = lst
lst += [3]
print("list :", lst, l_alias, "| same object:", lst is l_alias)

rule("4. the None sentinel defeats the mutable-default trap")

def bad(item, acc=[]):
    acc.append(item)
    return acc


def good(item, acc=None):
    if acc is None:
        acc = []
    acc.append(item)
    return acc


print("bad :", bad(1), bad(2), bad(3))
print("good:", good(1), good(2), good(3))

rule("5. shallow copy leaks, deep copy isolates")
inner = [1, 2]
shallow = copy.copy([inner])
deep = copy.deepcopy([inner])
inner.append(3)
print("shallow:", shallow)
print("deep   :", deep)
assert shallow == [[1, 2, 3]] and deep == [[1, 2]]

rule("6. attribute storage: dict versus __slots__")

class DictAttrs:
    def __init__(self, x, y):
        self.x, self.y = x, y


class Slotted:
    __slots__ = ("x", "y")

    def __init__(self, x, y):
        self.x, self.y = x, y


sizes = {}
for cls in (DictAttrs, Slotted):
    tracemalloc.start()
    objs = [cls(i, i) for i in range(100_000)]
    _, peak = tracemalloc.get_traced_memory()
    tracemalloc.stop()
    sizes[cls.__name__] = peak
    print(f"{cls.__name__:10s} peak {peak / 1024 / 1024:5.2f} MiB "
          f"instance {sys.getsizeof(objs[0])} bytes")
    del objs

saved = 100 * (1 - sizes["Slotted"] / sizes["DictAttrs"])
print(f"__slots__ saves {saved:.0f}% of instance memory")
assert saved > 20
print("\nAll object-model assertions passed.")

Line by line: the four asserts make the script fail loudly if a future CPython changes behaviour; int("256") versus int("257") isolates the small-int cache from literal folding; the two += blocks contrast in-place mutation with rebinding; bad versus good prints the exact trap and its fix; and tracemalloc measures real allocations where sys.getsizeof cannot — both classes report an identical 48-byte getsizeof, because the __dict__ is a separate object the function never accounts for.

Key takeaways and challenge

Challenge: write a Point class with __slots__, __eq__, and __hash__, then build a set of 100,000 points. Instrument it with tracemalloc and compare peak memory against the same set built from a dict-based class. Then deliberately mutate an x attribute of a point already in the set and show with in that the set can no longer find it.

Want one-to-one help building this kind of intuition? Ampersand Academy runs hands-on Python training.

Do Python variables hold values or references?

Names are bound to objects; assignment copies nothing. Writing b = a gives you a second name for the same object, so mutating through one name is visible through the other. Use a copy such as a[:] when you need independence.

When should I use is instead of double equals?

Use is only for singletons like None, True, False and NotImplemented. Identity for numbers and strings depends on caching and interning, so equality checks must use the equality operator.

Why does my function default argument keep accumulating values?

Default arguments are evaluated once when the function is defined, so a list default is shared across calls. Default to None and create the container inside the function instead.

What is the difference between a shallow copy and copy.deepcopy?

A shallow copy duplicates the outer container but shares the objects inside it, so mutating a nested list shows up in both. deepcopy recursively copies everything and is the safe choice for nested or shared structures.

Does __slots__ make Python objects smaller?

Yes. In a measured test of 100000 two-field instances, slots cut peak memory from 12.20 MiB to 8.38 MiB, about 31 percent, because each instance no longer carries its own attribute dictionary.

Exit mobile version