For this project, you will use a pre-trained deep neural network, SqueezeNet, which is lightweight and runs fast on CPUs. Run the code below to load a pre-trained SqueezeNet from the PyTorch official model zoo.
InΒ [2]:
# Test and set the device.if torch.cuda.is_available():Β Β Β device = βcuda:0βelse:Β Β Β device = βcpuβprint(βUseβ, device)Β # Download and load the pretrained SqueezeNet model.model = torchvision.models.squeezenet1_1(pretrained=True).to(device)Β # Disable the gradient computation with respect to model parameters.for param in model.parameters():Β Β Β param.requires_grad = False
Use cuda:0
DataΒΆ
For Task#1 and Task#2, use the images in folder Project1\images where the filenames are the corresponding class labels. For example, 182.png is an image of Border Terrier, which is class 182 in ImageNet dataset. Please refer to this Gist snippet for a complete list. The images are from ImageNet validation set, and so the pre-trained model has never βseenβ them.
For Task#3, you may use the images in folder Project1\style or any other images you like.
Helper FunctionsΒΆ
Most pre-trained models are trained on images that had been preprocessed by subtracting the per-color mean and dividing by the per-color standard deviation. Here are a few helper functions for performing and undoing this preprocessing.
InΒ [3]:
IMAGENET_MEAN = np.array([0.485, 0.456, 0.406])IMAGENET_STD = np.array([0.229, 0.224, 0.225])Β def preprocess(img, size=(224, 224)):Β Β Β transform = T.Compose([Β Β Β Β Β Β Β T.Resize(size),Β Β Β Β Β Β Β T.ToTensor(),Β Β Β Β Β Β Β T.Normalize(mean=IMAGENET_MEAN.tolist(),Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β std=IMAGENET_STD.tolist()),Β Β Β Β Β Β Β T.Lambda(lambda x: x[None]),Β Β Β ])Β Β Β return transform(img)Β def deprocess(img, should_rescale=True):Β Β Β transform = T.Compose([Β Β Β Β Β Β Β T.Lambda(lambda x: x[0]),Β Β Β Β Β Β Β T.Normalize(mean=[0, 0, 0], std=(1.0 / IMAGENET_STD).tolist()),Β Β Β Β Β Β Β T.Normalize(mean=(-IMAGENET_MEAN).tolist(), std=[1, 1, 1]),Β Β Β Β Β Β Β T.Lambda(rescale) if should_rescale else T.Lambda(lambda x: x),Β Β Β Β Β Β Β T.ToPILImage(),Β Β Β ])Β Β Β return transform(img)Β def rescale(x):Β Β Β low, high = x.min(), x.max()Β Β Β x_rescaled = (x β low) / (high β low)Β Β Β return x_rescaledΒ def blur_image(X, sigma=1):Β Β Β X_np = X.cpu().clone().numpy()Β Β Β X_np = gaussian_filter1d(X_np, sigma, axis=2)Β Β Β X_np = gaussian_filter1d(X_np, sigma, axis=3)Β Β Β X.copy_(torch.Tensor(X_np).type_as(X))Β Β Β return X
Task#1 Adversarial AttackΒΆ
The concept of βimage gradientsβ can be used to study the stability of a network. Consider a state-of-the-art deep neural network that generalizes well on an object recognition task. We expect such network to be robust to small perturbations to its input, because small perturbations cannot change the object category of an image. However, it was shown in the following paper[1] that by applying an imperceptible non-random perturbation to a test image, it is possible to arbitrarily change the networkβs prediction.
[1] Szegedy et al, βIntriguing properties of neural networksβ, ICLR 2014
Given an image and a target class, we can perform gradient ascent over the image to maximize the target class, stopping when the network classifies the image as the target class. While the perturbations seem negligible to humans, the network would classify the perturbed images wrongly.
Read the paper, and then implement the following function make_adversarial_attack to generate βfooling imagesβ. For each image in Project1/images with class label $c$, generate a fooling image that will be classified into class $ c-1-d $ where $d$ is the last digit of your student number. Save each fooling image into the folder Project1/fooling_images with the filename {true_class}_{target_class}.png. You may confirm (optional) that the fooling image 182_9.png in the folder will be wrongly classified as ostrich (class 9 in ImageNet dataset).
For the image 182.png, show the difference map between the original image and the fooling image, and save it as 182_1x_diff.png. Magnify the difference by 10 times and save the resulting map as 182_10x_diff.png.
InΒ [4]:
def make_adversarial_attack(X, target_y, model):Β Β Β βββΒ Β Β Generate a fooling image that is close to X, but that the model classifiesΒ Β Β as target_y.Β Β Β Β Inputs:Β Β Β β X: Input image; Tensor of shape (1, 3, 224, 224)Β Β Β β target_y: An integer in the range [0, 1000)Β Β Β β model: A pretrained CNNΒ Β Β Β Returns:Β Β Β β X_fooling: An image that is close to X, but that is classifed as target_yΒ Β Β by the model.Β Β Β βββΒ Β Β Β Β Β Β model.eval()Β Β Β Β Β Β Β # Initialize our fooling image to the input imageΒ Β Β X_fooling = X.clone().detach()Β Β Β X_fooling.requires_grad = TrueΒ Β Β Β # you may change the learning rate and max_iterΒ Β Β learning_rate = 1Β Β Β Β ##############################################################################Β Β Β # TODO: Generate a fooling image X_fooling that the model will classify asΒ Β #Β Β Β # the class target_y. You should perform gradient ascent on the score of the #Β Β Β # target class, stopping when the model is fooled.Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β #Β Β Β # When computing an update step, first normalize the gradient:Β Β Β Β Β Β Β Β Β Β Β Β Β Β #Β Β Β #Β Β dX = learning_rate * g / ||g||_2Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β #Β Β Β ##############################################################################Β Β Β loss = torch.nn.CrossEntropyLoss()Β Β Β # The coefficients are tried out by human for several hoursΒ Β Β # It is much better by using ones initialization ranther than randn initilizationΒ Β Β r = 0.001*torch.ones(1,3,224,224,device=device)Β Β Β r.requires_grad_(True)Β Β Β # training the fooling imageΒ Β Β for i in range(10000):Β Β Β Β Β Β Β X_fooling = X_fooling + rΒ Β Β Β Β Β Β output = model(X_fooling)Β Β Β Β Β Β Β output = torch.nn.functional.softmax(output,dim=1)Β Β Β Β Β Β Β # IPNN loss functionΒ Β Β Β Β Β Β cost =Β 0.00001*torch.norm(r,p=1) + loss(output,target_y)Β Β Β Β Β Β Β # BPΒ Β Β Β Β Β Β cost.backward(retain_graph=True)Β Β Β Β Β Β Β # Update the Β Β Β Β Β Β Β Β with torch.no_grad():Β Β Β Β Β Β Β Β Β Β Β r -= learning_rate*r.grad/torch.norm(r.grad,p=2)Β Β Β Β Β Β Β Β Β Β Β # Manually zero the gradients after updating weightsΒ Β Β Β Β Β Β Β Β Β Β r.grad.zero_()Β Β Β Β Β Β Β # Predict yΒ Β Β Β Β Β Β predict_y = torch.argmax(output)Β Β Β Β Β Β Β if predict_y == target_y:Β Β Β Β Β Β Β Β Β Β Β print(βInterations: %4dΒ Β Β Attack successedοΌβ %(i),βPredict_y:%dβ%(predict_y),βnorm(r,2):%fβ%(torch.norm(r,p=2).item()))Β Β Β Β Β Β Β Β Β Β Β breakΒ Β Β ##############################################################################Β Β Β #Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β END OF YOUR CODEΒ Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β #Β Β Β ##############################################################################Β Β Β X_fooling = X_fooling.detach()Β Β Β Β Β Β Β return X_fooling
InΒ [5]:
############################################################################### TODO: 1. Compute the fooling images for the images under `Project1/images`.##Β Β Β Β Β Β 2. Show the 4 related images of the image β182.pngβ: original image, ##Β Β Β Β Β Β Β Β Β fooling image, 182_1x_diff.png and2 182_10x_diff.png.Β Β Β Β Β Β Β Β Β Β Β Β Β ################################################################################ read imagesimport glob,osfiles = glob.glob(βimages\*.pngβ)# deal with imagesimages = []for file in files:Β Β Β print(file)Β Β Β input_image = PIL.Image.open(file)Β Β Β input_tensorΒ = preprocess(input_image).cuda()Β Β Β target_y = int(file[7:-4])-1-7 # A0206597UΒ Β Β print(target_y)Β Β Β target_y = torch.tensor([target_y]).cuda()Β Β Β output_tensor = make_adversarial_attack(input_tensor,target_y,model).cpu()Β Β Β image = deprocess(output_tensor)Β Β Β images.append(image)Β Β Β # save the imagesΒ Β Β image.save(os.path.join(βfooling_imagesβ,file[7:-4]+β_β+str(target_y.item())+β.pngβ)) Β # plot imagesoriginal_182 = PIL.Image.open(files[1])fig1 = plt.figure()plt.imshow(original_182)plt.title(βoriginal_182β)fig2 = plt.figure()fooling_182 = images[1]plt.imshow(fooling_182)plt.title(βfooling_182β)# diff imagestarget_y = int(files[1][7:-4])-1-7 # A0206597Utarget_y = torch.tensor([target_y]).cuda()input_tensorΒ = preprocess(original_182).cuda()output_tensor = make_adversarial_attack(input_tensor,target_y,model).cpu()difference_1x_182 = deprocess(input_tensor.cpu()-output_tensor,should_rescale=True)difference_10x_182 = deprocess((input_tensor.cpu()-output_tensor)*10,should_rescale=False)fig3 = plt.figure()plt.imshow(difference_1x_182)plt.title(βdifference_1x_182β)fig4 = plt.figure()plt.imshow(difference_10x_182)plt.title(βdifference_10x_182β)# save the diff imagesdifference_1x_182.save(β182_1x_diff.pngβ)difference_10x_182.save(β182_10x_diff.pngβ)###############################################################################Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β END OF YOUR CODEΒ Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β ###############################################################################
images\100.png
92
Interations:Β Β Β 7Β Β Β Attack successedοΌ Predict_y:92 norm(r,2):2.178698
images\182.png
174
Interations:Β Β Β 5Β Β Β Attack successedοΌ Predict_y:174 norm(r,2):2.717995
images\294.png
286
Interations:Β Β Β 4Β Β Β Attack successedοΌ Predict_y:286 norm(r,2):2.086930
images\366.png
358
Interations:Β Β Β 3Β Β Β Attack successedοΌ Predict_y:358 norm(r,2):2.122715
images\662.png
654
Interations:Β Β Β 5Β Β Β Attack successedοΌ Predict_y:654 norm(r,2):2.321405
images\85.png
77
Interations:Β Β 12Β Β Β Attack successedοΌ Predict_y:77 norm(r,2):2.496330
Interations:Β Β Β 5Β Β Β Attack successedοΌ Predict_y:174 norm(r,2):2.717996
Task#2 Class VisualizationΒΆ
By starting with a random noise image and performing gradient ascent on a target class, we can generate an image that the network will recognize as the target class. This idea was first presented in [2]; [3] extended this idea by suggesting several regularization techniques that can improve the quality of the generated image.
Concretely, let $I$ be an image and let $y$ be a target class. Let $s_y(I)$ be the score that a convolutional network assigns to the image $I$ for class $y$; note that these are raw unnormalized scores, not class probabilities. We wish to generate an image $I^*$ that achieves a high score for the class $y$ by solving the problem
$$ I^* = \arg\max_I (s_y(I) β R(I)) $$
where $R$ is a (possibly implicit) regularizer (note the sign of $R(I)$ in the argmax: we want to minimize this regularization term). We can solve this optimization problem using gradient ascent, computing gradients with respect to the generated image. We will use (explicit) L2 regularization of the form
$$ R(I) = \lambda \|I\|_2^2 $$
and implicit regularization as suggested by [3] by periodically blurring the generated image. We can solve this problem using gradient ascent on the generated image.
In the cell below, complete the implementation of the create_class_visualization function.
InΒ [6]:
def jitter(X, ox, oy):Β Β Β βββΒ Β Β Helper function to randomly jitter an image.Β Β Β Β Β InputsΒ Β Β β X: PyTorch Tensor of shape (N, C, H, W)Β Β Β β ox, oy: Integers giving number of pixels to jitter along W and H axesΒ Β Β Β Β Β Β Returns: A new PyTorch Tensor of shape (N, C, H, W)Β Β Β βββΒ Β Β if ox != 0:Β Β Β Β Β Β Β left = X[:, :, :, :-ox]Β Β Β Β Β Β Β right = X[:, :, :, -ox:]Β Β Β Β Β Β Β X = torch.cat([right, left], dim=3)Β Β Β if oy != 0:Β Β Β Β Β Β Β top = X[:, :, :-oy]Β Β Β Β Β Β Β bottom = X[:, :, -oy:]Β Β Β Β Β Β Β X = torch.cat([bottom, top], dim=2)Β Β Β return X
InΒ [7]:
def create_class_visualization(target_y, model, device, **kwargs):Β Β Β ββΒ Β Β Generate an image to maximize the score of target_y under a pretrained model.Β Β Β Β Β Β Β Inputs:Β Β Β β target_y: A list of two elements, where the first value is an integer in the range [0, 1000) giving the index of theΒ Β Β Β Β Β Β Β Β Β Β Β Β Β Β class, and the second value is the name of the class.Β Β Β β model: A pretrained CNN that will be used to generate the imageΒ Β Β β dtype: Torch datatype to use for computationsΒ Β Β Β Β Β Β Keyword arguments:Β Β Β β l2_reg: Strength of L2 regularization on the imageΒ Β Β β learning_rate: How big of a step to takeΒ Β Β β num_iterations: How many iterations to useΒ Β Β β blur_every: How often to blur the image as an implicit regularizerΒ Β Β β max_jitter: How much to gjitter the image as an implicit regularizerΒ Β Β β show_every: How often to show the intermediate resultΒ Β Β ββΒ Β Β model.to(device)Β Β Β l2_reg = kwargs.pop(βl2_regβ, 1e-3)Β Β Β learning_rate = kwargs.pop(βlearning_rateβ, 25)Β Β Β num_iterations = kwargs.pop(βnum_iterationsβ, 100)Β Β Β blur_every = kwargs.pop(βblur_everyβ, 10)Β Β Β max_jitter = kwargs.pop(βmax_jitterβ, 16)Β Β Β show_every = kwargs.pop(βshow_everyβ, 25)Β Β Β Β Β Β Β # Randomly initialize the image as a PyTorch Tensor, and make it requires gradient.Β Β Β img = torch.randn(1, 3, 224, 224).mul_(1.0).to(device).requires_grad_()Β Β Β Β for t in range(num_iterations):Β Β Β Β Β Β Β # Randomly jitter the image a bit; this gives slightly nicer resultsΒ Β Β Β Β Β Β ox, oy = random.randint(0, max_jitter), random.randint(0, max_jitter)Β Β Β Β Β Β Β img.data.copy_(jitter(img.data, ox, oy))Β Β Β Β Β Β Β Β ########################################################################Β Β Β Β Β Β Β # TODO: Use the model to compute the gradient of the score for theΒ Β Β Β #Β Β Β Β Β Β Β # class target_y with respect to the pixels of the image, and make aΒ Β #Β Β Β Β Β Β Β # gradient step on the image using the learning rate. Donβt forget the #Β Β Β Β Β Β Β # L2 regularization term!Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β #Β Β Β Β Β Β Β # Be very careful about the signs of elements in your code.Β Β Β Β Β Β Β Β Β Β Β #Β Β Β Β Β Β Β ########################################################################Β Β Β Β Β Β Β scores = model(img)Β Β Β Β Β Β Β loss = l2_reg*torch.norm(img,p=2)-2*scores[0,target_y[0]]Β Β Β Β Β Β Β loss.backward()Β Β Β Β Β Β Β # Update imgΒ Β Β Β Β Β Β with torch.no_grad():Β Β Β Β Β Β Β Β Β Β Β img -= learning_rate*img.gradΒ Β Β Β Β Β Β Β Β Β Β # Manually zero the gradients after updating weightsΒ Β Β Β Β Β Β Β Β Β Β img.grad.zero_()Β Β Β Β Β Β Β ########################################################################Β Β Β Β Β Β Β #Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β END OF YOUR CODEΒ Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β #Β Β Β Β Β Β Β ########################################################################Β Β Β Β Β Β Β Β # Undo the random jitterΒ Β Β Β Β Β Β img.data.copy_(jitter(img.data, -ox, -oy))Β Β Β Β Β Β Β Β # As regularizer, clamp and periodically blur the imageΒ Β Β Β Β Β Β for c in range(3):Β Β Β Β Β Β Β Β Β Β Β lo = float(-IMAGENET_MEAN[c] / IMAGENET_STD[c])Β Β Β Β Β Β Β Β Β Β Β hi = float((1.0 β IMAGENET_MEAN[c]) / IMAGENET_STD[c])Β Β Β Β Β Β Β Β Β Β Β img.data[:, c].clamp_(min=lo, max=hi)Β Β Β Β Β Β Β if t % blur_every == 0:Β Β Β Β Β Β Β Β Β Β Β blur_image(img.data, sigma=0.5)Β Β Β Β Β Β Β Β # Periodically show the imageΒ Β Β Β Β Β Β if t == 0 or (t + 1) % show_every == 0 or t == num_iterations β 1:Β Β Β Β Β Β Β Β Β Β Β plt.imshow(deprocess(img.data.clone().cpu()))Β Β Β Β Β Β Β Β Β Β Β class_name = target_y[1]Β Β Β Β Β Β Β Β Β Β Β plt.title(β%s\nIteration %d / %dβ % (class_name, t + 1, num_iterations))Β Β Β Β Β Β Β Β Β Β Β plt.gcf().set_size_inches(4, 4)Β Β Β Β Β Β Β Β Β Β Β plt.axis(βoffβ)Β Β Β Β Β Β Β Β Β Β Β plt.show()Β Β Β Β return deprocess(img.data.cpu())
Once you have completed the implementation in the cell above, run the following cell to generate an image of a Tarantula:
InΒ [8]:
target_y = [76, βTarantulaβ]# target_y = 366 # Gorillaimport pdbout = create_class_visualization(target_y, model, device)
Task#3 Style TransferΒΆ
Another task which is closely related to image gradients is style transfer which has become a βcoolβ application in deep learning for computer vision applications. You need to study and implement the style transfer technique presented in the following paper [4] where the general idea is to take two images (a content image and a style image), and produce a new image that reflects the content of one but the artistic βstyleβ of the other.
Below is an example.
Compute the lossΒΆ
To perform style transfer, you will need to first formulate a special loss function that matches the content and style of each respective image in the feature space, and then perform gradient descent on the pixels of the image itself.
The loss function contains two parts: content loss and style loss. Read the paper [4] for details about the losses and implement them below.
InΒ [9]:
def content_loss(content_weight, content_current, content_original):Β Β Β βββΒ Β Β Compute the content loss for style transfer.Β Β Β Β Β Β Β Inputs:Β Β Β β content_weight: Scalar giving the weighting for the content loss.Β Β Β β content_current: features of the current image; this is a PyTorch Tensor of shapeΒ Β Β Β Β (1, C_l, H_l, W_l).Β Β Β β content_target: features of the content image, Tensor with shape (1, C_l, H_l, W_l).Β Β Β Β Β Β Β Returns:Β Β Β β scalar content lossΒ Β Β βββΒ Β Β Β Β Β Β ##############################################################################Β Β Β # TODO: Implement content loss functionΒ Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β #Β Β Β # Note: It should not be very much code (less than 10 lines)Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β #Β Β Β ##############################################################################Β Β Β loss = content_weight*torch.sum((content_current-content_original)**2)Β Β Β return lossΒ Β Β ##############################################################################Β Β Β #Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β END OF YOUR CODEΒ Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β #Β Β Β ##############################################################################Β Β def gram_matrix(features):Β Β Β βββΒ Β Β Compute the normalized Gram matrix from features.Β Β Β The Gram matrix will be used to compute style loss.Β Β Β Β Β Β Β Inputs:Β Β Β β features: PyTorch Tensor of shape (N, C, H, W) giving features forΒ Β Β Β Β a batch of N images.Β Β Β Β Β Β Β Returns:Β Β Β β gram: PyTorch Tensor of shape (N, C, C) giving theΒ Β Β Β Β normalized Gram matrices for the N input images.Β Β Β βββΒ Β Β Β Β Β Β ##############################################################################Β Β Β # TODO: Implement the normalized Gram matrix compuation functionΒ Β Β Β Β Β Β Β Β Β Β Β #Β Β Β # Note: It should not be very much code (less than 10 lines)Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β #Β Β Β ##############################################################################Β Β Β N,C,H,W = features.size()Β Β Β features = torch.reshape(features,(N,C,-1))Β Β Β gram = torch.bmm(features,features.permute(0,2,1))/float(C*H*W)Β Β Β return gramΒ Β Β ##############################################################################Β Β Β #Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β END OF YOUR CODEΒ Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β #Β Β Β ##############################################################################Β Β def style_loss(feats, style_layers, style_targets, style_weights):Β Β Β βββΒ Β Β Computes the style loss at a set of layers.Β Β Β Β Β Β Β Inputs:Β Β Β β feats: list of the features at every layer of the current image.Β Β Β β style_layers: List of layer indices into feats giving the layers to include in theΒ Β Β Β Β style loss.Β Β Β β style_targets: List of the same length as style_layers, where style_targets[i] isΒ Β Β Β Β a PyTorch Variable giving the Gram matrix of the source style image computed atΒ Β Β Β Β layer style_layers[i].Β Β Β β style_weights: List of the same length as style_layers, where style_weights[i]Β Β Β Β Β is a scalar giving the weight for the style loss at layer style_layers[i].Β Β Β Β Β Β Β Β Β Returns:Β Β Β β style_loss: A PyTorch Tensor holding a scalar giving the style loss.Β Β Β βββ Β Β Β Β Β Β Β ##############################################################################Β Β Β # TODO: Implement style loss functionΒ Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β #Β Β Β # Note: It should not be very much code (less than 10 lines)Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β #Β Β Β ##############################################################################Β Β Β loss = 0Β Β Β for i,layer in enumerate(style_layers):Β Β Β Β Β Β Β cur_gram = gram_matrix(feats[layer])Β Β Β Β Β Β Β loss += style_weights[i]*torch.sum((cur_gram-style_targets[i])**2)/4Β Β Β return lossΒ Β Β ##############################################################################Β Β Β #Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β END OF YOUR CODEΒ Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β #Β Β Β ##############################################################################
Putting them togetherΒΆ
With these loss functions, you can now build your style transfer model. Implement the function below to perform style transfer. To test the model, you can use the content and style images that we have provided in Project1/style, or improvise using any image you like. Please save your output images in the Project1/style folder.
Design and carry out some experiments (on your own!) to analyse how the choice of layers and the weights will influence the output image. Write down your observations and analysis in the Markdown cell provided below.
InΒ [146]:
def style_transfer(content_image, style_image, content_layer, content_weight,Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β style_layers, style_weights, max_iter):Β Β Β βββΒ Β Β Run style transfer!Β Β Β You may first resize the image to a small size for fast computation.Β Β Β Β Β Β Β Inputs:Β Β Β β content_image: filename of content imageΒ Β Β β style_image: filename of style imageΒ Β Β β content_layer: an index indicating which layer to use for content lossΒ Β Β β content_weight: weighting on content lossΒ Β Β β style_layers: list of indices indicating which layers to use for style lossΒ Β Β β style_weights: list of weights to use for each layer in style_layersΒ Β Β β max_iter: max iterations of gradient updatesΒ Β Β Β Β Β Β Returns:Β Β Β β output_image: an image with content from the content_image and Β Β Β Β style from the style imageΒ Β Β βββΒ Β Β ##############################################################################Β Β Β # TODO: Implement the function for style transfer.Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β #Β Β Β ##############################################################################Β Β Β def extract_features(data,model):Β Β Β Β Β Β Β feats = []Β Β Β Β Β Β Β prev_feat = dataΒ Β Β Β Β Β Β for module in model.features:Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β next_feat = module(prev_feat)Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β feats.append(next_feat)Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β prev_feat = next_featΒ Β Β Β Β Β Β return featsΒ Β Β # remove white noiseΒ Β Β def variance_loss(image):Β Β Β Β Β Β Β loss = 0Β Β Β Β Β Β Β # rowΒ Β Β Β Β Β Β loss += torch.sum( ( image[:,:,1:,:] β image[:,:,:-1,:] )**2 ) Β Β Β Β Β Β Β Β # columnΒ Β Β Β Β Β Β loss += torch.sum( ( image[:,:,:,1:] β image[:,:,:,:-1] )**2 )Β Β Β Β Β Β Β return lossΒ Β Β # Content imageΒ Β Β content_img = preprocess(PIL.Image.open(content_image)).to(device)Β Β Β feats = extract_features(content_img,model)Β Β Β content_target = feats[content_layer].clone()Β Β Β # Style imageΒ Β Β style_img = preprocess(PIL.Image.open(style_image)).to(device)Β Β Β feats = extract_features(style_img,model)Β Β Β style_target = []Β Β Β for layer in style_layers:Β Β Β Β Β Β Β style_target.append(gram_matrix(feats[layer].clone()))Β Β Β # Transfer imageΒ Β Β tran_img = content_img.clone().requires_grad_()Β Β Β # TrainingΒ Β Β optimizer = torch.optim.Adam([tran_img],lr=1.0)Β Β Β # Plot input imagesΒ Β Β f, axarr = plt.subplots(1,2)Β Β Β axarr[0].axis(βoffβ)Β Β Β axarr[1].axis(βoffβ)Β Β Β axarr[0].set_title(βContent Source Img.β)Β Β Β axarr[1].set_title(βStyle Source Img.β)Β Β Β axarr[0].imshow(deprocess(content_img.cpu()))Β Β Β axarr[1].imshow(deprocess(style_img.cpu()))Β Β Β plt.show()Β Β Β plt.figure()Β Β Β for i in range(max_iter):Β Β Β Β Β Β Β # Enhancing contrast ratioΒ Β Β Β Β Β Β tran_img.data.clamp_(-1.5, 1.5)Β Β Β Β Β Β Β # Init gradΒ Β Β Β Β Β Β optimizer.zero_grad()Β Β Β Β Β Β Β feats = extract_features(tran_img, model)Β Β Β Β Β Β Β # Compute lossΒ Β Β Β Β Β Β c_loss = content_loss(content_weight, feats[content_layer], content_target)Β Β Β Β Β Β Β s_loss = style_loss(feats, style_layers, style_target, style_weights)Β Β Β Β Β Β Β v_loss = variance_loss(tran_img)Β Β Β Β Β Β Β total_loss = 0.2*c_loss + s_loss + 0.1*v_lossΒ Β Β Β Β Β Β # BPΒ Β Β Β Β Β Β total_loss.backward()Β Β Β Β Β Β Β optimizer.step()Β Β Β print(βIteration {}β.format(i))Β Β Β plt.axis(βoffβ)Β Β Β plt.imshow(deprocess(tran_img.data.cpu()))Β Β Β deprocess(tran_img.data.cpu()).save(βstyle/engineering-the_scream.pngβ)Β Β Β plt.show()Β Β Β ##############################################################################Β Β Β #Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β END OF YOUR CODEΒ Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β #Β Β Β ##############################################################################
InΒ [147]:
############################################################################### TODO: 1. Choose one pair of images under βProject1/styleβ, and finish theΒ # #Β Β Β Β Β Β Β Β Β neural style transfer task by calling the style_transfer function.##Β Β Β Β Β Β 2. Show the 3 related images: content image, style image and theΒ Β Β Β # #Β Β Β Β Β Β Β Β Β generated style-transferred image.Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β ################################################################################ Content layer comaparing params1 = {Β Β Β βcontent_imageβ : βstyle/engineering.jpgβ,Β Β Β βstyle_imageβ : βstyle/the_scream.jpgβ,Β Β Β βcontent_layerβ : 3,Β Β Β βcontent_weightβ : 0.1, Β Β Β Β βstyle_layersβ : (1, 3, 4, 6, 7),Β Β Β βstyle_weightsβ : (20000, 10000,500, 12, 1),Β Β Β βmax_iterβ : 100}style_transfer(**params1)###############################################################################Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β END OF YOUR CODEΒ Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β Β ###############################################################################
Iteration 99
As the content layer index increases, the transmitted image will look more like βstyle_imageβ. In addition, the ratio of $ \ alpha / \ beta $ also controls the similarity with βcontent_imageβ or βstyle_imageβ. The larger the loss caused by the style layer, the more it looks like a style image.



