Transfer Learning Image Classification: Cats vs. Dogs using Xception
This project demonstrates transfer learning using a pre-trained Xception model from TensorFlow/Keras to classify images as either cats or dogs. Transfer learning is a technique where a model trained on one large dataset (ImageNet) is reused and fine-tuned for a different but related task, dramatically reducing training time and improving performance on smaller datasets.
Key Concepts:
- Transfer Learning: Leverages pre-trained weights from Xception (trained on ImageNet) rather than training from scratch
- Binary Classification: Classifies images into two classes (cats or dogs) using sigmoid activation
- Data Augmentation: Uses ImageDataGenerator for preprocessing and normalization
- Custom Functional Model: Builds a hybrid model combining the pre-trained Xception base with custom dense layers
- Early Stopping: Implements a custom callback (MyCLRuleMonitor) to halt training when validation performance meets criteria
- Notebook:
11_TransferLearningCatAndDogsXception.ipynb - Dataset:
cats_and_dogs.zip(contains train/ and validation/ folders) - Model Architecture: Pre-trained Xception base + custom dense layers (128 → 128 → 1)
- Input Size: 128 × 128 × 3 (RGB images)
- Output: Binary classification (sigmoid, probability of dog = 1, cat = 0)
Python Version: 3.8+ (tested with 3.11/3.12/3.13)
Dependencies:
pip install --upgrade pip
pip install tensorflow numpy pandas pillow matplotlib scikit-learnOr install in one command:
pip install tensorflow>=2.10 numpy pandas pillow matplotlib scikit-learn- Extract the dataset:
import shutil
shutil.unpack_archive('cats_and_dogs.zip', 'cats_and_dogs')- Expected folder structure:
cats_and_dogs/
├── train/
│ ├── cats/
│ │ ├── image1.jpg
│ │ ├── image2.jpg
│ │ └── ...
│ └── dogs/
│ ├── image1.jpg
│ ├── image2.jpg
│ └── ...
└── validation/
├── cats/
│ └── ...
└── dogs/
└── ...
jupyter notebook 11_TransferLearningCatAndDogsXception.ipynbOr open directly in VS Code and run cells sequentially.
Step 1: Import Libraries
import pandas as pd
import numpy as np
import tensorflow as tfStep 2: Data Preprocessing
- Load train/validation data with ImageDataGenerator (rescale 1.0/255.0)
- Target size: 128×128 for efficient inference
- Class mode: binary (cats=0, dogs=1)
Step 3: Load Pre-trained Xception Model
xception = tf.keras.applications.xception.Xception(include_top=False)
# Freeze existing weights to preserve learned features
for layer in xception.layers:
layer.trainable = FalseStep 4: Build Custom Functional Model
- Input: 128×128×3 RGB image
- Xception base (frozen)
- Flatten → Dense(128, relu) → Dense(128, relu) → Dense(1, sigmoid)
Step 5: Compile and Train
model.compile(optimizer='adam', loss='binary_crossentropy', metrics=['accuracy'])
model.fit(train_image, validation_data=test_image, epochs=100, callbacks=[MyCLRuleMonitor(0.9)])Step 6: Predict on New Images
image = tf.keras.preprocessing.image.load_img('your_image.jpg', target_size=(128,128))
image_arr = tf.keras.preprocessing.image.img_to_array(image)
np_img_arr = np.expand_dims(image_arr, axis=0)
probability = model.predict(np_img_arr)
# If probability[0][0] >= 0.5 → Dog; else → CatInput (128, 128, 3)
↓
Xception (pre-trained, frozen)
↓
Flatten
↓
Dense(128, relu) [h1]
↓
Dense(128, relu) [h2]
↓
Dense(1, sigmoid) [output]
↓
Output (probability: 0=cat, 1=dog)
Cause: Prediction image size differs from training target size (128×128).
Fix: Always use the same target size:
image = tf.keras.preprocessing.image.load_img('image.jpg', target_size=(128, 128))Cause: The MyCLRuleMonitor callback compares numpy arrays or tensors directly without converting to scalar.
Fix: Update the callback to safely convert logs to scalars:
class MyCLRuleMonitor(tf.keras.callbacks.Callback):
def __init__(self, CL, metric_name='accuracy'):
super().__init__()
self.CL = float(CL)
self.metric_name = metric_name
def _to_scalar(self, value):
if value is None:
return None
try:
if isinstance(value, tf.Tensor):
value = value.numpy()
arr = np.asarray(value)
return float(arr.item() if arr.size == 1 else arr.mean())
except:
return None
def on_epoch_end(self, epoch, logs=None):
if logs is None:
return
train_val = self._to_scalar(logs.get(self.metric_name))
val_val = self._to_scalar(logs.get(f'val_{self.metric_name}'))
if train_val is not None and val_val is not None:
if (val_val > train_val) and (val_val >= self.CL):
self.model.stop_training = TrueCause: cats_and_dogs.zip not in the working directory.
Fix:
- Ensure
cats_and_dogs.zipis in the notebook's directory - Run extraction cell explicitly:
import shutil, os
if not os.path.exists('cats_and_dogs'):
shutil.unpack_archive('cats_and_dogs.zip', 'cats_and_dogs')
print('Dataset extracted successfully')
else:
print('Dataset already exists')Cause: Batch size too large or insufficient GPU/RAM.
Fix: Reduce batch size in flow_from_directory:
train_image = train_image_data.flow_from_directory(..., batch_size=16) # reduce from 20
test_image = test_image_data.flow_from_directory(..., batch_size=16)Cause:
- Pre-trained weights are frozen but dataset is very different
- Insufficient training epochs
- Learning rate too high/low
Fix:
- Unfreeze some top layers of Xception and fine-tune:
for layer in xception.layers[-20:]: # unfreeze last 20 layers
layer.trainable = True
model.compile(optimizer=tf.keras.optimizers.Adam(learning_rate=1e-5), ...)- Increase epochs or use learning rate scheduling
- Add Data Augmentation to the ImageDataGenerator for better generalization:
train_image_data = tf.keras.preprocessing.image.ImageDataGenerator(
rescale=1.0/255.0,
rotation_range=20,
horizontal_flip=True,
zoom_range=0.2
)- Use a Better Pooling Layer instead of Flatten for robustness:
x = tf.keras.layers.GlobalAveragePooling2D()(xception(input_layer))- Add Dropout to reduce overfitting:
x2 = tf.keras.layers.Dropout(0.5)(x1)- Save and Load Model:
model.save('cats_dogs_xception.keras')
loaded_model = tf.keras.models.load_model('cats_dogs_xception.keras')- Fine-tune Xception by unfreezing upper layers for better domain adaptation
- Initial Training: Keep Xception frozen (faster, good baseline)
- Fine-tuning: Unfreeze top 20-50 layers with very low learning rate (1e-5)
- Early Stopping: Use the callback with CL=0.85-0.95 to avoid overfitting
- Batch Size: Start with 16-32 for stability
- Epochs: 50-100 epochs usually sufficient with pre-trained model
With the pre-trained Xception model:
- Training accuracy: 95%+
- Validation accuracy: 85-92% (depending on dataset quality)
- Training time: 5-15 minutes per epoch (on CPU) or 1-3 minutes (on GPU)
- Xception base model: ~81 MB
- Full model after training: ~85 MB
- VRAM required: 2-4 GB (GPU) or 8+ GB (CPU recommended)
- Xception Paper - Chollet, 2016
- TensorFlow Transfer Learning Guide
- Keras ImageDataGenerator
- Keras Callbacks
No specific license provided. Add appropriate license as needed (MIT, Apache 2.0, GPL, etc.).
Created for Simplylearn - AIML Course (Module 4) as part of Deep Learning and Transfer Learning examples.
Questions or Issues? Check the troubleshooting section above or refer to the cell comments in the notebook.