Python & R Programming — Complete Revision Guide

C-DAC Kharghar · PG-DBDA February 2026 · Beginner → Advanced

36h Theory 44h Lab 40h Self Learning 14 Python Sessions 4 R Sessions

This guide covers every session of the PG-DBDA Python & R module at C-DAC Kharghar. Each topic is explained with real-world Indian analogies, illustrated with SVG diagrams, and backed by working, exam-ready code. Whether you are encountering programming for the first time or brushing up before the lab exam, read each session from the top — concepts build on each other deliberately.

14
Python Sessions
4
R Sessions
18
SVG Diagrams
120+
Code Snippets
📋 Evaluation Pattern
ComponentWeightageWhat it tests
Theory Exam40%Concepts, definitions, tracing code, explaining data structures
Lab Exam40%Writing programs, debugging, output prediction
Internal20%Assignments, quizzes, mini-projects, class participation
"Programs must be written for people to read, and only incidentally for machines to execute." — Harold Abelson
🗺️ Learning Roadmap

Sessions 1–2 · Python Foundations

Variables · Data Types · if/else · for/while loops · Strings · Input/Output

Sessions 3–4 · Core Data Structures

Dictionaries · Lists · Comprehensions · Functional tools (map, filter, zip)

Sessions 5–7 · Functions & Tuples

Lambda · Recursion · Modules · Pickling · Tuples as fixed records

Sessions 8–10 · Advanced Python

OOP · Generators · Decorators · Regex · Exception Handling · Logging

Sessions 11–13 · Data Science Toolkit

NumPy · Pandas · Matplotlib · Seaborn · Plotly · BeautifulSoup · Pillow

Session 14 · Python + Databases

SQLite · MySQL · Pandas + SQL · CRUD operations · Parameterised queries

Sessions 15–18 · R Programming

Basics · Vectors · Data Frames · tidyverse · ggplot2 · R Markdown

⚡ Jump to Any Topic

🔀 Control Flow

if/else, for, while, break, continue — decision and repetition

Session 1

📝 Strings

Slicing, methods, palindromes, pangrams, f-strings

Session 2

📚 Dictionaries

Key-value, comprehension, ROT-13 cipher lab

Session 3

⚙️ Functions

Lambda, map, recursion, *args, closures, pickling

Sessions 5–6

🏗️ OOP

Classes, inheritance, generators, decorators, regex

Sessions 8–9

🐼 Pandas & NumPy

DataFrames, arrays, groupby, diamonds dataset

Session 11

📊 R Basics

Vectors, assignment, arithmetic, RStudio exploration

Session 15

🌊 tidyverse

dplyr pipes, ggplot2 charts, data reshaping

Session 17

Session 1 · Python Basics & Control Flow

Variables · Data Types · Operators · if/else · for · while · break/continue/pass

2h Theory4h Lab2h Self-Learning
🎯 The Core Analogy

Python is like giving instructions to a very literal-minded cook. You write a recipe (program), the cook follows it exactly — no assumptions, no shortcuts. An if-else is: "If the milk is sour, skip it; otherwise use it." A for loop is: "For each student in the class roll call, say their name." The cook never improvises — that precision is Python's strength.

How Python runs your code

Python is an interpreted language — unlike C or Java which compile to machine code first, Python reads your file and executes it line by line. Understanding this journey helps you understand import errors, bytecode caching, and performance characteristics.

your_code.py plain text file Python Interpreter reads line-by-line Bytecode (.pyc) compiled form Python VM executes → output No separate compile step — python3 myfile.py does everything automatically
Beginner Data Types
Python's Built-in Data Types
TypeExampleC-DAC Real-world UsageType Check
introll = 2601Roll number, batch count, port numbertype(roll)
floatcgpa = 8.75CGPA, salary (₹), ML model accuracyisinstance(x, float)
strname = "Priya"Names, SQL queries, file paths, JSON keystype(x) == str
boolpassed = TruePass/Fail flags, ON/OFF switches, DB flagsisinstance(x, bool)
NoneTyperesult = NoneMissing data, NULL from DB, unset variableresult is None
complexz = 3+4jSignal processing, DSP, FFTz.real, z.imag
Python · Variables & Types
# Python is DYNAMICALLY TYPED — no type declaration needed
name       = "Priya Nair"       # str
roll       = 2601              # int
cgpa       = 8.92              # float
is_placed  = True              # bool
dept       = None              # NoneType — not yet assigned

# Python infers the type at runtime
print(type(cgpa))     # <class 'float'>

# Multiple assignment on one line
x, y, z = 10, 20, 30

# Swap without a temp variable (very Pythonic!)
x, y = y, x           # x=20, y=10 — tuple packing/unpacking

# Type conversion (casting)
int("42")      # → 42        (string to int)
str(3.14)     # → "3.14"    (float to string)
float(7)      # → 7.0       (int to float)
bool(0)       # → False     (0, "", [], None are all Falsy)
bool("CDAC") # → True      (any non-empty string is Truthy)
Beginner Control Flow — if / elif / else

Control flow tells Python which path to take through your code. Python uses indentation (4 spaces) to define code blocks — no curly braces like Java or C. Wrong indentation is always a SyntaxError.

marks = input() program starts if marks ≥ 90? first condition check Yes Grade O Outstanding No elif marks ≥ 60? second check Yes Grade B First class No else Grade F — Fail
Python · Grade Calculator (C-DAC Style)
marks = int(input("Enter marks (0-100): "))

if marks >= 90:
    print("O — Outstanding")
elif marks >= 75:
    print("A — Distinction")
elif marks >= 60:
    print("B — First Class")
elif marks >= 40:
    print("C — Pass")
else:
    print("F — Fail. Please reappear.")

# One-liner ternary — for simple True/False choices
result = "Pass" if marks >= 40 else "Fail"
print(f"Result: {result}")
BeginnerIntermediate Loops — for & while
for loop Use when you KNOW the count for i in range(10): for s in class_list: for char in "CDAC": Analogy: class roll call — 30 students while loop Use when CONDITION drives exit while balance > 0: while not found: while user_input != 'q': Analogy: wait at IRCTC until seat confirmed
Python · Lab Q1 & Q2 — Factorials
# Q1: for loop — print 0! to 10!
f, n = 1, 0
for _ in range(11):
    print(f"{n}! = {f}")
    n += 1
    f *= n
# Output: 0!=1, 1!=1, 2!=2, ... 10!=3,628,800

# Q2: while loop — all factorials under 2 billion
f, n = 1, 0
while f < 2_000_000_000:      # underscore separators for readability
    print(f"{n}! = {f:,}")    # :, adds comma formatting to numbers
    n += 1
    f *= n
break · continue · pass
KeywordEffectAnalogyCommon Use
breakExit entire loop nowEmergency stop on a machineFound what you searched for
continueSkip this iteration, go to nextSkip a corrupted track on SpotifySkip invalid data rows
passDo nothing — placeholderEmpty room stub in a blueprintSkeleton class/function
🧪 What does print(type(True)) output?

Session 2 · Strings

Text manipulation — the single most critical skill for data cleaning and preprocessing

2h Theory2h Lab2h Self-Learning
🎯 The Core Analogy

A string is a Mumbai local train with coaches numbered from 0. Coach 0 is the first, coach -1 is the last (counting from the back). You can board at any coach (s[3]), or take a stretch of coaches (s[2:7]), or walk the entire train backwards (s[::-1]). The train is immutable — you can not rebuild a coach in place. s[0] = 'X' gives a TypeError; you must produce a new string.

s = "C-DAC Kharghar" C - D A C ⎵ K 0 1 2 3 4 5 6 … -14 -13 -11 -10 s[0]→'C' s[4]→'C' s[-1]→'r' s[0:4]→'C-DA' s[6:]→'Kharghar' s[::-1] reverses entire string
Python · String Operations — Complete Reference
s = "C-DAC Kharghar Navi Mumbai"

# SLICING — [start:stop:step]  (stop is exclusive!)
print(s[0:5])        # 'C-DAC'
print(s[6:])         # 'Kharghar Navi Mumbai'
print(s[::-1])       # reverse the whole string
print(s[6:14:2])     # every 2nd char from 6 to 14

# IMPORTANT METHODS
print(s.upper())              # 'C-DAC KHARGHAR NAVI MUMBAI'
print(s.lower())              # 'c-dac kharghar navi mumbai'
print(s.replace('Navi','New'))# replace substring
print(s.split(' '))           # ['C-DAC', 'Kharghar', 'Navi', 'Mumbai']
print(', '.join(['a','b','c'])) # 'a, b, c'
print(s.strip())              # remove leading/trailing whitespace
print(s.count('a'))          # count occurrences of 'a'
print(s.find('Kharghar'))    # index position or -1 if not found
print(s.startswith('C-DAC'))# True
print(s.endswith('Mumbai'))  # True

# F-STRINGS — modern, fast, readable string formatting
batch, cgpa = 2026, 8.9247
print(f"PG-DBDA {batch} — C-DAC Kharghar")
print(f"CGPA: {cgpa:.2f}")        # 'CGPA: 8.92' (2 decimal places)
print(f"Rank: {1:04d}")           # 'Rank: 0001' (zero-padded)
print(f"Total: ₹{75000:,}")      # 'Total: ₹75,000'
Intermediate Lab Q3 & Q4 — Palindrome & Pangram
Python · Lab Solutions
import re

# Q3: Phrase palindrome checker
def is_palindrome(phrase):
    """Strip non-letters, lowercase, compare to reverse."""
    clean = re.sub(r'[^a-zA-Z]', '', phrase).lower()
    return clean == clean[::-1]

tests = ["Was it a rat I saw?", "Step on no pets", "Dammit, I'm mad!", "Hello CDAC"]
for t in tests:
    print(f"{'✓' if is_palindrome(t) else '✗'}  {t!r}")

# Q4: Pangram checker (contains all 26 English letters)
def is_pangram(sentence):
    letters = set(sentence.lower())
    return set('abcdefghijklmnopqrstuvwxyz').issubset(letters)

print(is_pangram("The quick brown fox jumps over the lazy dog"))  # True
print(is_pangram("Hello World"))                               # False

Session 3 · Dictionaries

Key-Value stores — the data structure powering Big Data, JSON, and every web API

2h Theory2h Lab
🎯 The Core Analogy

A Python dict is exactly like your phone's contact book. You search by name (key) and get the phone number (value) instantly — you do not scroll through every contact. Lookup is O(1) whether you have 10 or 10 million entries. This is the same principle behind Hadoop's shuffle-sort phase, Redis, and MongoDB document stores.

KEYS (must be immutable) VALUES (any Python object) "name" "Priya Nair" "roll" 2601 (integer) "marks" {"Python":88,"R":79} (dict!) O(1) average lookup via internal hash table · keys immutable · values = any object · insertion-ordered (Python 3.7+)
Python · Dictionary Essentials + Lab Q1 ROT-13
# CREATE
student = {
    "name": "Priya Nair",
    "roll": 2601,
    "marks": {"Python": 88, "R": 79},  # nested dict!
    "batch": "PGDBDA-Feb2026"
}

# ACCESS — two ways
print(student["name"])              # KeyError if key missing
print(student.get("dept", "N/A"))  # safe — returns default

# MODIFY & ADD
student["marks"]["Java"] = 83     # add to nested dict
student["graduated"] = False       # add new top-level key

# DELETE
del student["roll"]
popped = student.pop("batch")    # remove and return value

# ITERATION
for key, val in student.items():
    print(f"  {key}: {val}")

# DICT COMPREHENSION — like list comp but for dicts
squares = {x: x**2 for x in range(1, 6)}  # {1:1, 2:4, 3:9, 4:16, 5:25}

# ── ROT-13 CIPHER LAB Q1 ──────────────────────────────────────
def rot13(text):
    upper = 'ABCDEFGHIJKLMNOPQRSTUVWXYZ'
    lower = upper.lower()
    table = str.maketrans(
        upper + lower,
        upper[13:] + upper[:13] + lower[13:] + lower[:13]
    )
    return text.translate(table)

secret = "Pnrfne pvcure? V zhpu cersre Pnrfne fnynq!"
print(rot13(secret))   # Caesar cipher? I much prefer Caesar salad!
print(rot13(rot13(secret)) == secret)  # True — self-inverse!

Session 4 · Lists

Python's most versatile data structure — ordered, mutable, and infinitely flexible

2h Theory2h Lab
🎯 The Core Analogy

A list is like a Zomato order cart. You can add items (append), remove one (remove), look at what's at the top ([-1]), merge two carts (+), and even create a mini-cart of just the first three items ([:3]). Everything is ordered by position, and nothing is fixed — it is completely mutable.

Python · Complete List Operations Reference
scores = [88, 72, 95, 61, 84, 91]

--- ADDING ---
scores.append(78)            # add to end
scores.insert(0, 100)         # insert at index 0
scores.extend([55, 67])      # merge another list
scores += [82]               # shorthand extend

--- REMOVING ---
scores.remove(61)            # remove first occurrence of value 61
scores.pop()                 # remove & return last element
scores.pop(2)                # remove & return element at index 2
del scores[0:2]              # delete a slice
scores.clear()               # empty the entire list

--- SEARCHING & STATS ---
print(max(scores), min(scores), sum(scores))
print(scores.index(95))      # position of 95
print(scores.count(88))      # how many times 88 appears
print(95 in scores)          # True — membership test

--- SORTING ---
scores.sort()                # ascending, modifies IN-PLACE
scores.sort(reverse=True)   # descending, in-place
ranked = sorted(scores)     # returns NEW sorted list, original unchanged

--- LIST COMPREHENSION (Pythonic & fast!) ---
passed   = [s for s in scores if s >= 40]
doubled  = [s * 2 for s in scores]
grades   = ['P' if s>=40 else 'F' for s in scores]

--- CRITICAL: Copy vs Reference ---
a = [1, 2, 3]
b = a           # SAME list! b IS a — dangerous!
c = a.copy()   # true independent copy
d = a[:]        # slice copy — also independent
e = list(a)    # another way to copy
# Test: b.append(99) → a also has 99. c.append(99) → a unchanged.
⚠️

Most common beginner mistake: b = a does NOT copy a list — it makes b point to the same list in memory as a. If you do b.append(99), you will find 99 in a as well. Always use a.copy() or a[:] when you need an independent copy.

Sessions 5–6 · Functions & Modules

Lambda · Recursion · *args/**kwargs · Modules · Pickling — reusable code building blocks

4h Theory4h Lab4h Self-Learning
🎯 The Core Analogy

A function is a mixer-grinder. You put ingredients in (parameters), it processes, you get output (return value). Once you own it, use it anytime (reusability). A lambda is a small hand blender — quick, one-time use, no need to give it a shelf. A module is the entire appliance catalogue from a brand — grouped, importable, ready to use.

def calculate_gpa ( m1, m2, m3 ): keyword function name parameters (inputs) gpa = (1/3)*m1 + (1/2)*m2 + (1/4)*m3 ← function body (indented 4 spaces) return gpa ← sends output back to caller
Python · Function Types — All Variants
# 1. Named function with default argument
def greet(name, batch="PGDBDA-2026"):
    return f"Welcome {name} to C-DAC Kharghar — {batch}!"

greet("Priya")                # uses default batch
greet("Raj", "AI-ML-2026")   # overrides default

# 2. *args — variable positional arguments (packed as tuple)
def total_marks(*subjects):
    print(f"Got {len(subjects)} subjects")
    return sum(subjects)

total_marks(88, 79, 91, 83)    # 341

# 3. **kwargs — variable keyword arguments (packed as dict)
def student_info(**details):
    for k, v in details.items():
        print(f"  {k}: {v}")

student_info(name="Zara", roll=2610, cgpa=9.1)

# 4. Lambda — anonymous, one-liner, no return keyword
square    = lambda x: x**2
is_even   = lambda n: n % 2 == 0

# Lambda with sorted() — sort students by marks (highest first)
students = [("Priya",88), ("Raj",74), ("Zara",95), ("Dev",81)]
ranked = sorted(students, key=lambda s: s[1], reverse=True)
# [('Zara', 95), ('Priya', 88), ('Dev', 81), ('Raj', 74)]

# 5. Recursive function (Lab Q2 — factorial)
def factorial(n):
    if n == 0: return 1          # BASE CASE — must exist!
    return n * factorial(n - 1)  # recursive case

for i in range(1, 11):
    print(f"{i:2d}  {factorial(i):>10,}")
Call Stack for factorial(4) — how recursion unwinds factorial(4) = 4 × ? factorial(3) = 3 × ? factorial(2) = 2 × ? factorial(0) → 1 (BASE CASE stops recursion) ← unwinds: 4 × 6 = 24 ← unwinds: 3 × 2 = 6 ← unwinds: 2 × 1 = 2
Python · Pickling — Object Serialisation
import pickle

model_data = {
    "batch": "PGDBDA-2026",
    "students": ["Priya", "Raj", "Zara"],
    "accuracy": 0.94
}

# Serialize (pack) to disk — wb = write binary
with open("batch.pkl", "wb") as f:
    pickle.dump(model_data, f)

# Deserialize (unpack) — rb = read binary
with open("batch.pkl", "rb") as f:
    loaded = pickle.load(f)

print(loaded["students"])   # ['Priya', 'Raj', 'Zara']
# Used heavily in ML to save/load trained scikit-learn models!

Session 7 · Tuples — Deep Dive

Immutable sequences — faster, safer, hashable. The fixed-schema record of Python.

2h Theory2h Lab
🎯 The Core Analogy

A tuple is a sealed Aadhaar card. Once issued, the data is fixed — you cannot change your date of birth in it. A list is an editable form. Use tuples for fixed facts: GPS coordinates (18.92, 72.83), student record (2601, "Priya", 88.5), RGB colour (255, 0, 0), DB row. These facts should never change in your program.

Use Tuple when…

Data is fixed and should never change: DB rows, coordinates, RGB values, function returning multiple values, dictionary keys (tuples are hashable — lists are not)

Use List when…

Data grows or shrinks: exam scores accumulating, a shopping cart, log entries, any collection that changes during the program's lifetime

Python · Tuple Operations — All Lab Exercises
# Create
coords  = (18.9220, 72.8347)   # lat/lon near C-DAC Kharghar area
rgb     = (255, 140, 0)          # dark orange
single  = (42,)                  # MUST have trailing comma for single-element!
t       = (10, 20, 30, 20, 10)

# Lab Q1: Find repeated items
repeated = [x for x in set(t) if t.count(x) > 1]
print(repeated)     # [10, 20]

# Lab Q2: Sort list of tuples by float element
items = [('item1', '12.20'), ('item2', '15.10'), ('item3', '24.5')]
items_sorted = sorted(items, key=lambda x: float(x[1]), reverse=True)
# [('item3','24.5'), ('item2','15.10'), ('item1','12.20')]

# Lab Q3: Count elements until a tuple is found
mixed = [10, 20, 30, (40, 50), 60]
count = next(i for i, x in enumerate(mixed) if isinstance(x, tuple))
print(count)        # 3

# Lab Q4: Element-wise sum using zip()
t1, t2, t3 = (1,2,3,4), (3,5,2,1), (2,2,3,1)
result = tuple(a+b+c for a,b,c in zip(t1, t2, t3))
print(result)       # (6, 9, 8, 6)

# Tuple unpacking — very Pythonic
lat, lon = coords
x, y, *rest = (1, 2, 3, 4, 5)   # rest = [3,4,5]

Sessions 8–9 · OOP, Generators, Decorators & Regex

The professional Python toolkit — how every production system is actually built

4h Theory4h Lab5h Self-Learning
🎯 OOP — The Core Analogy

A Class is the blueprint; an Object is the actual building. C-DAC has one blueprint called Student with attributes (name, roll, marks) and methods (calculate_gpa, display). Every enrolled student — Priya, Raj, Zara — is an object created from that blueprint. Same blueprint, thousands of instances, each with their own data.

CLASS: Student — blueprint — Attributes: name, roll, marks Methods: __init__() calculate_gpa() has_failed() __str__() instance instance s1: Student name = "Priya" roll = 2601 marks = {Py:88,R:79} s2: Student name = "Raj" roll = 2602 s3: Student name = "Zara" roll = 2610 marks = {Py:95,R:88}
Python · Complete Student Class — Lab Q1 & Q2
class Student:
    course = "PG-DBDA"               # class variable — shared by ALL objects

    def __init__(self, roll, name, marks_dict):
        # Instance variables — unique to EACH object
        self.roll  = roll
        self.name  = name
        self.marks = marks_dict        # {subject: [list of 5 marks]}

    def __str__(self):               # called by print(student)
        return f"[{self.roll}] {self.name} — {self.course}"

    def calculate_gpa(self, m1, m2, m3):
        # GPA formula from syllabus
        return (1/3)*m1 + (1/2)*m2 + (1/4)*m3

    def has_failed(self):
        # any() short-circuits — stops at first True found
        return any(
            m < 40
            for mlist in self.marks.values()
            for m in mlist
        )

# Inheritance — Faculty extends Person
class Person:
    def __init__(self, name): self.name = name
    def speak(self): return "..."

class Faculty(Person):                # inherits Person
    def speak(self):                  # method OVERRIDING
        return f"Prof. {self.name}"

# Polymorphism — same method call, different behaviour
people = [Faculty("Dr. Mehta"), Person("Visitor")]
for p in people: print(p.speak())
Advanced Generators & Decorators
ℹ️

Why Generators matter for Big Data: A list of 10 million rows takes ~400MB RAM. A generator produces one item at a time, using barely any memory — which is exactly how Spark processes huge datasets in chunks. yield pauses the function and returns one value; the next call resumes from where it stopped.

Python · Generators + Decorators — Lab Q3
# Generator — memory-efficient sequence
class NumberSeries:
    def even_numbers(self, limit):
        n = 0
        while n <= limit:
            yield n      # pauses here; next() call resumes
            n += 2

ns = NumberSeries()
for num in ns.even_numbers(20):
    print(num, end=" ")   # 0 2 4 6 8 10 12 14 16 18 20

# Decorator — wraps a function to add behaviour WITHOUT changing it
# Analogy: a SmartWatch wraps your wrist — your wrist is unchanged
import functools, time

def log_method_call(func):
    @functools.wraps(func)           # preserves original function name
    def wrapper(*args, **kwargs):
        print(f"→ Calling {func.__name__}({args[1:]})")
        start = time.time()
        result = func(*args, **kwargs)
        print(f"← Done in {time.time()-start:.4f}s")
        return result
    return wrapper

class Student:
    @log_method_call                   # apply decorator with @
    def calculate_gpa(self, m1, m2, m3):
        return (1/3)*m1+(1/2)*m2+(1/4)*m3

s = Student()
s.calculate_gpa(80, 75, 90)
# → Calling calculate_gpa((80, 75, 90))
# ← Done in 0.0001s

Session 10 · Exception Handling & Logging

Writing code that survives the real world — graceful failure over catastrophic crashes

2h Theory2h Lab
🎯 The Core Analogy

Exception handling is an emergency protocol at a hospital. Things you don't expect will go wrong. You don't prevent every scenario — you have trained responders (except clauses), cleanup teams (finally blocks), and post-incident reports (logging). Without this, one unexpected user input can bring down an entire production system running FRAS or Digi-Exam.

Python · Full Exception Handling + Custom Exceptions + Logging
try:
    marks = int(input("Enter marks: "))   # ValueError if "abc"
    avg   = 500 / marks                    # ZeroDivisionError if 0
    print(f"Average: {avg:.1f}")

except ValueError:
    print("Please enter a valid integer.")

except ZeroDivisionError:
    print("Marks cannot be zero.")

except Exception as e:
    print(f"Unexpected: {type(e).__name__}: {e}")

else:
    print("Calculation successful.")   # runs ONLY if no exception

finally:
    print("Done.")                       # ALWAYS runs (cleanup)

# CUSTOM EXCEPTION — for domain-specific validation
class InvalidMarksError(Exception):
    def __init__(self, val):
        super().__init__(f"Marks {val} out of range (0–100)")

def validate_marks(m):
    if not 0 <= m <= 100:
        raise InvalidMarksError(m)    # explicitly raise
    return m

# LOGGING — structured, persistent, production-grade
import logging
logging.basicConfig(
    level   = logging.DEBUG,
    format  = '%(asctime)s [%(levelname)s] %(message)s',
    handlers = [
        logging.FileHandler('cdac_app.log'),
        logging.StreamHandler()          # also print to console
    ]
)
logging.info("FRAS v2.0 started — C-DAC Kharghar")
logging.warning("DB connection latency > 500ms")
logging.error("Face recognition model load failed")
logging.critical("Disk full — attendance records at risk!")
Logging Levels — Severity Scale (Low → High)
LevelNumberWhen to useExample
DEBUG10Detailed tracing during development"Processing student ID 2601"
INFO20Normal operation milestones"API server started on port 8000"
WARNING30Unexpected but recoverable"Disk usage at 85%"
ERROR40Serious failure, function aborted"Cannot connect to PostgreSQL"
CRITICAL50System about to die"Out of memory — killing process"

Session 11 · Pandas & NumPy

The data science powerhouses — arrays, DataFrames, wrangling, and the diamonds dataset lab

2h Theory6h Lab2h Self-Learning
🎯 The Core Analogy

NumPy is a scientific calculator; Pandas is Excel inside Python. NumPy works on entire arrays at once — no Python loops, C-speed math. Pandas gives you labelled rows and columns (like an Excel sheet) with SQL-like groupby, merge, and pivot. Together they turn weeks of manual spreadsheet work into a few lines of code.

NumPy Homogeneous arrays (all same type) Vectorised math — no Python loops arr[arr > 85] — boolean mask Foundation of Pandas, Scikit-learn, TF Best for: number crunching, images Pandas Labelled rows + columns (DataFrame) Mixed types allowed per column SQL-like: groupby, merge, pivot Built on NumPy under the hood Best for: tabular data, CSV, SQL output
Python · NumPy + Pandas — Lab Q1–Q4 (Diamonds Dataset)
import numpy as np
import pandas as pd

═══ NumPy Essentials ═══════════════════════════════════════
marks = np.array([88, 92, 76, 95, 61, 84])

# Vectorised — all elements at once
print(marks + 5)                   # add 5 to every element
print(marks[marks > 85])           # [88, 92, 95] — boolean indexing
print(marks.mean(), marks.std())   # statistics
print(np.percentile(marks, [25,50,75]))

# Image array (Lab Q2)
from PIL import Image
img = Image.open('photo.jpg')
arr = np.array(img)               # (H, W, 3) array
restored = Image.fromarray(arr)

═══ Pandas — Diamonds Dataset (Lab Q4) ═════════════════════
df = pd.read_csv('diamonds.csv')

# Print first 6 rows
print(df.head(6))

# Mean of each numeric column
print(df.mean(numeric_only=True))

# Count / min / max price for each cut
print(df.groupby('cut')['price'].agg(['count','min','max']))

# Concise summary
df.info()            # dtypes, non-null counts
df.describe()        # mean, std, quartiles

# Count duplicate rows
print(df.duplicated().sum())

# Data cleaning
df.dropna(inplace=True)         # drop rows with NaN
df.drop_duplicates(inplace=True)

# Add computed column
df['price_per_carat'] = df['price'] / df['carat']

# Filter
premium = df[(df['cut']=='Ideal') & (df['price']>5000)]

Session 12 · Visualisation & Web Scraping

matplotlib · seaborn · plotly · ggplot · BeautifulSoup4 — making data speak visually

2h Theory4h Lab2h Self-Learning
🎯 The Core Analogy

Visualisation is the translation layer between numbers and human understanding. A table of 10,000 rows tells you nothing at a glance. A heatmap of correlations tells you everything in 3 seconds. The best data scientists spend more time on charts than on models — because a wrong model caught by a good plot saves weeks of wasted computation. BeautifulSoup is how you collect that data from the web in the first place.

matplotlib foundation — full control, any plot seaborn statistics + beauty plotly interactive, web-ready plotnine / ggplot grammar of graphics pandas .plot() quick EDA shortcut All high-level libraries wrap matplotlib under the hood — understanding matplotlib = understanding all of them
Beginner Library Selection Guide
Which Library for Which Job?
LibraryBest ChartsStrengthWeakness
matplotlibAny — line, bar, scatter, hist, pie, 3DAbsolute control, publication-ready, subplotsVerbose code for simple charts
seabornheatmap, boxplot, violin, pairplot, distplotStatistical defaults, beautiful out-of-boxLess flexible than matplotlib
plotlyAny + 3D, choropleth, sunburst, candlestickInteractive (hover, zoom, pan), web embedHeavy file size, slower render
plotnineR-style ggplot grammar in PythonLayered aesthetic mapping, familiar to R usersSlower, smaller community
df.plot()Basic line, bar, hist, scatter, boxZero extra import — just call on DataFrameLimited customisation

Part 1 — matplotlib: Anatomy & Core Plots

matplotlib works on a Figure → Axes → Plot hierarchy. The Figure is the entire canvas. The Axes is one chart area inside it (you can have multiple). Every element — title, labels, ticks, gridlines, legend — is individually controllable.

Figure (the whole canvas — fig = plt.figure()) Y-axis label X-axis label Chart Title (ax.set_title()) Series A Series B Axes (ax) fig, ax = plt.subplots() ax.set_title("...") ax.set_xlabel("...") ax.set_ylabel("...") ax.legend() ax.grid(True) ax.set_xlim(0, 10) plt.tight_layout() plt.savefig("out.png") plt.show()
Python · matplotlib — Complete Chart Gallery
import matplotlib.pyplot as plt
import numpy as np

═══ 1. Line Chart ═══════════════════════════════════════════
months = ['Jan','Feb','Mar','Apr','May','Jun']
sales  = [120, 145, 132, 178, 190, 210]

fig, ax = plt.subplots(figsize=(9, 4))
ax.plot(months, sales, marker='o', color='#b34a22',
        linewidth=2, label='Monthly Sales')
ax.fill_between(months, sales, alpha=0.1, color='#b34a22')
ax.set_title('C-DAC Sales Dashboard')
ax.set_xlabel('Month'); ax.set_ylabel('Units Sold')
ax.legend(); ax.grid(True, alpha=0.3)
plt.tight_layout(); plt.show()

═══ 2. Bar Chart ════════════════════════════════════════════
subjects = ['Python', 'R', 'Stats', 'Java', 'Linux']
avg_marks = [82, 76, 88, 71, 85]
colors   = ['#b34a22' if m>=80 else '#d4c8b0' for m in avg_marks]

fig, ax = plt.subplots(figsize=(8, 4))
bars = ax.bar(subjects, avg_marks, color=colors, edgecolor='white')
ax.bar_label(bars, fmt='%d')                # add value labels on bars
ax.axhline(80, color='#1a6e5c', linestyle='--', label='Pass Line')
ax.set_title('PG-DBDA Batch Average Marks')
ax.set_ylim(0, 100); ax.legend()
plt.tight_layout(); plt.show()

═══ 3. Histogram ════════════════════════════════════════════
marks_data = np.random.normal(72, 12, 200)  # simulate marks distribution
fig, ax = plt.subplots(figsize=(8, 4))
ax.hist(marks_data, bins=20, color='#1a6e5c', edgecolor='white', alpha=0.8)
ax.axvline(marks_data.mean(), color='#b34a22', linestyle='--', label=f'Mean={marks_data.mean():.1f}')
ax.set_title('Marks Distribution'); ax.legend()
plt.tight_layout(); plt.show()

═══ 4. Subplot Grid — show multiple charts at once ══════════
fig, axes = plt.subplots(1, 3, figsize=(14, 4))

axes[0].plot([1,2,3], [4,5,6]); axes[0].set_title('Line')
axes[1].bar(['A','B','C'], [3,7,5]); axes[1].set_title('Bar')
axes[2].scatter([1,2,3], [3,1,4]); axes[2].set_title('Scatter')

plt.suptitle('Multi-panel Dashboard')
plt.tight_layout(); plt.show()

═══ 5. Save figure to file ══════════════════════════════════
fig.savefig('report_chart.png', dpi=150, bbox_inches='tight')
fig.savefig('report_chart.pdf')          # vector format for reports

Part 2 — seaborn: Statistical Visualisation

seaborn is built on top of matplotlib but adds a higher-level interface designed specifically for statistical graphics. It understands Pandas DataFrames natively — pass column names as strings and seaborn does the grouping, colouring, and labelling for you.

Python · seaborn — Statistical Charts + Lab Q2 Full Solution
import seaborn as sns
import matplotlib.pyplot as plt
import pandas as pd
import numpy as np

═══ Seaborn Themes — set once, applies everywhere ═══════════
sns.set_theme(style='whitegrid', palette='muted')
# Other styles: darkgrid, white, ticks, dark

═══ 1. Heatmap — Lab Q2 Step 1 & 2 ════════════════════════
df = pd.read_csv('data.csv')
corr = df.corr(numeric_only=True)

plt.figure(figsize=(10, 8))
sns.heatmap(
    corr,
    annot     = True,         # show values in cells
    fmt       = '.2f',         # 2 decimal places
    cmap      = 'RdYlGn',      # Red-Yellow-Green diverging
    center    = 0,             # 0 = white centre
    linewidths= 0.5,           # cell borders
    square    = True,          # force square cells
    vmin=-1, vmax=1            # fix scale -1 to +1
)
plt.title('Correlation Matrix — PG-DBDA Dataset', pad=14)
plt.tight_layout(); plt.show()

═══ 2. Find highest correlated pair (Lab Q2 Step 3) ════════
mask = np.triu(np.ones_like(corr, dtype=bool))   # upper triangle
col1, col2 = corr.mask(mask).stack().idxmax()
print(f"Highest correlation: {col1} ↔ {col2}")

═══ 3. Scatter plot (Lab Q2 Step 4) ════════════════════════
fig, ax = plt.subplots(figsize=(8, 6))
sns.scatterplot(data=df, x=col1, y=col2,
                alpha=0.6, color='#b34a22', ax=ax)
sns.regplot(data=df, x=col1, y=col2, scatter=False,
            color='#1a6e5c', ax=ax, label='Trend line')
ax.set_title(f'{col1} vs {col2} (r={corr.loc[col1,col2]:.2f})')
ax.legend()
plt.tight_layout(); plt.show()

═══ 4. Box Plot — compare distributions by category ════════
tips = sns.load_dataset('tips')           # built-in sample dataset
plt.figure(figsize=(9, 5))
sns.boxplot(data=tips, x='day', y='total_bill', hue='sex')
plt.title('Bill Distribution by Day and Gender')
plt.show()

═══ 5. Violin Plot — like boxplot but shows distribution shape
plt.figure(figsize=(9, 5))
sns.violinplot(data=tips, x='day', y='total_bill',
               hue='sex', split=True, palette='muted')
plt.title('Bill Distribution — Violin (shows density shape)')
plt.show()

═══ 6. Pair Plot — all columns vs all columns at once ══════
iris = sns.load_dataset('iris')
sns.pairplot(iris, hue='species', diag_kind='kde')
plt.suptitle('Iris Dataset — Pairwise Relationships', y=1.02)
plt.show()

═══ 7. Count Plot — categorical frequency bar ═══════════════
sns.countplot(data=tips, x='day', order=['Thur','Fri','Sat','Sun'])
plt.title('Number of Customers per Day')
plt.show()
💡

seaborn vs matplotlib — key rule: Start with seaborn for speed. When you need to customise something seaborn can't do, use ax.set_*() methods on the underlying matplotlib Axes object that seaborn returns. They work together seamlessly.

Part 3 — plotly: Interactive Visualisation

plotly produces charts that live in a browser — you can hover, zoom, pan, click to filter, and export as PNG. This is what you embed in Dash dashboards, Flask web apps, and HTML reports. The plotly.express module gives you the same one-liner simplicity as seaborn, but interactive.

Python · plotly.express — Interactive Charts (Lab Q2 Step 5)
import plotly.express as px
import plotly.graph_objects as go
import pandas as pd

═══ 1. Interactive Scatter with trendline (Lab Q2) ══════════
df = pd.read_csv('data.csv')
fig = px.scatter(
    df, x=col1, y=col2,
    trendline    = "ols",          # OLS regression line
    color        = 'category',     # colour by category column
    hover_data   = ['id'],          # show extra info on hover
    title        = f"{col1} vs {col2} (Interactive)"
)
fig.update_layout(template='plotly_white')
fig.show()                          # opens in browser
fig.write_html("scatter.html")      # save as shareable HTML

═══ 2. Interactive Bar Chart ════════════════════════════════
fig = px.bar(
    df.groupby('cut')['price'].mean().reset_index(),
    x='cut', y='price',
    color='cut', text_auto='.0f',
    title='Average Diamond Price by Cut'
)
fig.show()

═══ 3. Sunburst — hierarchical breakdown ════════════════════
fig = px.sunburst(
    px.data.tips(),
    path=['day', 'time', 'sex'],
    values='total_bill',
    title='Bill Breakdown (Day → Time → Gender)'
)
fig.show()

═══ 4. 3D Scatter ═══════════════════════════════════════════
fig = px.scatter_3d(
    px.data.iris(),
    x='sepal_length', y='sepal_width', z='petal_length',
    color='species', symbol='species',
    title='Iris Dataset — 3D View'
)
fig.show()

═══ 5. Animated Chart — add time dimension ══════════════════
fig = px.scatter(
    px.data.gapminder(),
    x='gdpPercap', y='lifeExp',
    size='pop', color='continent',
    hover_name='country', animation_frame='year',
    log_x=True, title='Gapminder — GDP vs Life Expectancy'
)
fig.show()

Part 4 — ggplot / plotnine: Grammar of Graphics

The Grammar of Graphics is a principled way to describe charts. Every chart is described as: data + aesthetic mapping + geometric layer + optional scales/facets. This is the native language of R's ggplot2, and Python's plotnine brings this exact grammar to Python. Understanding it makes you fluent in both Python and R visualisation.

data DataFrame + aes() x=col, y=col, color= + geom_*() point / bar / line / box + facet / theme split panels, styling Chart rendered output ggplot(data, aes(x='carat', y='price', color='cut')) + geom_point() + facet_wrap('~cut')
Python · plotnine (ggplot) — Grammar of Graphics
# pip install plotnine
from plotnine import *
import pandas as pd
import seaborn as sns

diamonds = sns.load_dataset('diamonds')

═══ Basic scatter ════════════════════════════════════════════
p = (ggplot(diamonds, aes(x='carat', y='price', color='cut'))
     + geom_point(alpha=0.3, size=0.5)
     + labs(title='Diamond Price vs Carat',
            x='Carat Weight', y='Price (USD)')
     + theme_minimal()
)
print(p)

═══ Faceted chart — one panel per cut category ═══════════════
p2 = (ggplot(diamonds.sample(2000), aes(x='carat', y='price'))
      + geom_point(color='#b34a22', alpha=0.5, size=1)
      + geom_smooth(method='lm', color='#1a6e5c')
      + facet_wrap('~cut', ncol=3)            # one panel per cut
      + labs(title='Price vs Carat by Cut Quality')
      + theme_bw()
)
print(p2)

═══ Bar + coord_flip ════════════════════════════════════════
p3 = (ggplot(diamonds, aes(x='cut', fill='cut'))
      + geom_bar()
      + coord_flip()                           # horizontal bars
      + theme_classic()
      + labs(title='Diamond Cut Counts')
)
print(p3)

Part 5 — pandas .plot(): Quick EDA

Python · pandas Built-in Plotting — Fastest EDA
import pandas as pd
import matplotlib.pyplot as plt

df = pd.read_csv('diamonds.csv')

# Line — index vs value
df['price'][:200].plot(title='First 200 Prices'); plt.show()

# Histogram of numeric columns
df.hist(figsize=(12,8), bins=30, color='#b34a22'); plt.show()

# Box plot by group
df.boxplot(column='price', by='cut'); plt.show()

# Grouped bar chart
df.groupby('cut')['price'].mean().plot(
    kind='bar', color='#1a6e5c', rot=0
)
plt.title('Mean Price by Cut'); plt.show()

Part 6 — BeautifulSoup: Web Scraping

Web scraping is the process of programmatically extracting data from websites. The workflow is always: fetch the HTML → parse it → locate the elements → extract the data → clean and store. BeautifulSoup handles steps 2–4; requests handles step 1; Pandas handles step 5.

Website HTML source requests.get() HTTP GET → raw HTML BeautifulSoup parse → find_all → extract Clean Data list / dict DataFrame CSV / DB / viz Always check robots.txt · Respect rate limits · Never scrape personal or sensitive data
Python · BeautifulSoup — Lab Q1 + Advanced Patterns
import requests
from bs4 import BeautifulSoup
import pandas as pd
import time

═══ Basic Scrape (Lab Q1) ════════════════════════════════════
# Set a header so you look like a real browser
headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0)"}

resp = requests.get("https://books.toscrape.com", headers=headers)
print(resp.status_code)    # 200 = success, 403 = blocked, 404 = not found

soup = BeautifulSoup(resp.text, 'html.parser')

# Find all book titles
titles = [t.find('a')['title'] for t in soup.find_all('h3')]
prices = [p.text for p in soup.find_all('p', class_='price_color')]
ratings= [a['class'][1] for a in soup.find_all('p', class_='star-rating')]

for t, p, r in zip(titles[:5], prices[:5], ratings[:5]):
    print(f"{t[:35]:35s} | {p} | {r} stars")

═══ CSS Selector approach (alternative to find_all) ══════════
# soup.select() uses CSS selectors — very powerful!
book_articles = soup.select('article.product_pod')

books = []
for article in book_articles:
    books.append({
        'title' : article.select_one('h3 a')['title'],
        'price' : article.select_one('.price_color').text,
        'rating': article.select_one('p.star-rating')['class'][1],
        'in_stock': 'In stock' in article.select_one('.availability').text
    })

df_books = pd.DataFrame(books)
print(df_books.head())
df_books.to_csv('books_scraped.csv', index=False)

═══ Multi-page Scraping (paginate through all 50 pages) ══════
all_books = []
base_url = "https://books.toscrape.com/catalogue/page-{}.html"

for page in range(1, 6):           # scrape first 5 pages
    url  = base_url.format(page)
    resp = requests.get(url, headers=headers)
    soup = BeautifulSoup(resp.text, 'html.parser')

    for article in soup.select('article.product_pod'):
        all_books.append({
            'title': article.select_one('h3 a')['title'],
            'price': article.select_one('.price_color').text,
            'page':  page
        })

    time.sleep(1)                   # BE POLITE — 1 second delay between pages!
    print(f"Page {page} done — {len(all_books)} books so far")

df_all = pd.DataFrame(all_books)
print(f"Total scraped: {len(df_all)} books")

═══ Key Parsing Methods Reference ═══════════════════════════
# soup.find('tag')           — first matching tag
# soup.find_all('tag')       — all matching tags (returns list)
# soup.find('tag', class_='x') — tag with specific class
# soup.find('tag', id='x')   — tag with specific id
# soup.select('div.card h2') — CSS selector (most flexible)
# soup.select_one('...')     — CSS selector, first match only
# tag.text / tag.get_text()  — extract text content
# tag['href'] / tag['class'] — extract attribute value
# tag.get('attr', default)   — safe attribute access
⚠️

Web Scraping Ethics & Legality: (1) Always check robots.txt — if the site disallows scraping, do not proceed. (2) Add time.sleep(1) between requests — do not hammer servers. (3) Never scrape personal data (names, emails, phone numbers) without consent. (4) For heavy scraping, use an API if one exists — it is faster, more reliable, and officially permitted. Legal violations can result in banning or litigation.

BeautifulSoup Parsing Methods — Quick Reference
MethodWhat it returnsExample
soup.find('h1')First <h1> tag object (or None)soup.find('h1').text
soup.find_all('a')List of all <a> tags[a['href'] for a in soup.find_all('a')]
soup.select('div.card')List via CSS selectorsoup.select('.product h3 a')
soup.select_one('p.price')First match via CSS selectorsoup.select_one('.price').text
tag.textAll text content of tagp_tag.text.strip()
tag['attr']Attribute valuea_tag['href'], img_tag['src']
tag.get('attr', '')Attribute value with default (safe)a_tag.get('class', [])
tag.parentParent elementspan.parent.find('h2')
tag.next_siblingNext sibling elementNavigate DOM tree
🧪 You want a chart where you can hover over data points to see exact values and zoom in/out — which Python library is the best choice?

Session 13 · Pillow, Audio & Virtual Environments

Image processing · Audio files · Isolated Python environments for production

2h Theory2h Lab3h Self-Learning
Python · Pillow + NumPy + VEnv Commands
═══ Pillow — Image Lab ═════════════════════════════════════
from PIL import Image
import numpy as np

img = Image.open('photo.jpg')
print(img.size, img.mode)      # (1920, 1080) RGB

# Convert to NumPy array (Lab Q2)
arr = np.array(img)
print(arr.shape)                # (1080, 1920, 3) — H×W×channels
print(arr[0, 0])                # [255 255 255] — top-left pixel RGB

# Back to image
restored = Image.fromarray(arr)
restored.save('output.jpg')

# Transforms
img.resize((800, 600))           # new size
img.rotate(90)                   # rotate 90°
img.convert('L')                 # RGB → Grayscale
img.crop((100,100,500,400))      # (left, top, right, bottom)

═══ Virtual Environment — Essential Commands ════════════════
# Create
python3 -m venv myproject_env

# Activate
source myproject_env/bin/activate      # Linux / Mac
myproject_env\Scripts\activate         # Windows PowerShell

# Install packages (only inside this env)
pip install fastapi uvicorn pandas numpy pillow

# Freeze for sharing / deployment
pip freeze > requirements.txt

# Recreate on another machine / production server
pip install -r requirements.txt

# Deactivate
deactivate

Session 14 · Python + Databases

Connecting Python to SQLite, MySQL, and PostgreSQL — the data pipeline bridge

2h Theory2h Lab4h Self-Learning
🎯 The Core Analogy

Python is the waiter; the database is the kitchen. The waiter (Python) takes your order (SQL query), walks it to the kitchen (DB engine), brings the food (result set), and serves it to you. The connection is the bridge. Always close it when done — like clocking out after shift — or you leak resources.

Python · SQLite + Pandas Integration (NEVER use string formatting for queries!)
import sqlite3, pandas as pd

# Connect (creates file if not exists)
conn = sqlite3.connect('cdac_attendance.db')
cur  = conn.cursor()

# Create table
cur.execute("""
    CREATE TABLE IF NOT EXISTS students (
        id     INTEGER PRIMARY KEY AUTOINCREMENT,
        name   TEXT    NOT NULL,
        roll   TEXT    UNIQUE,
        cgpa   REAL    DEFAULT 0.0,
        batch  TEXT
    )
""")

# Insert — use ? placeholders (NEVER string format — SQL injection!)
batch_data = [
    ("Priya Nair",  "CDAC2601", 8.9, "PGDBDA-26"),
    ("Raj Kumar",   "CDAC2602", 7.4, "PGDBDA-26"),
    ("Zara Sheikh", "CDAC2610", 9.1, "PGDBDA-26"),
]
cur.executemany(
    "INSERT OR IGNORE INTO students(name,roll,cgpa,batch) VALUES(?,?,?,?)",
    batch_data
)
conn.commit()

# Query directly into Pandas DataFrame — best practice!
df = pd.read_sql("SELECT * FROM students WHERE cgpa > 8", conn)
print(df)

# Always close
conn.close()

═══ MySQL Connection ═══════════════════════════════════════
# pip install mysql-connector-python
import mysql.connector
conn = mysql.connector.connect(
    host="localhost", user="root",
    password="cdac@123", database="pgdbda"
)
# Identical cursor/execute/commit/close API as SQLite

═══ PostgreSQL Connection ══════════════════════════════════
# pip install psycopg2-binary
import psycopg2
conn = psycopg2.connect(
    host="localhost", port=5432,
    dbname="fras_db", user="cdac", password="cdac@123"
)
# Use %s instead of ? for parameterised queries in psycopg2
🧪 Why should you NEVER use f-strings to build SQL queries like f"SELECT * FROM users WHERE name='{user_input}'"?

Session 15 · Introduction to R

Statistical computing — built for data by statisticians, beloved by researchers worldwide

2h Theory2h Lab2h Self-Learning
🎯 The Core Analogy

If Python is a Swiss Army knife, R is a surgeon's precision scalpel. Python does everything — web, ML, automation, APIs, databases. R is purpose-built for statistical analysis, hypothesis testing, and publication-quality charts. A data scientist at a pharmaceutical company, IIM, or ISRO research lab will almost certainly use R.

Why R over Python for stats?

10,000+ CRAN packages for every statistical method. Built-in distributions and hypothesis tests. ggplot2 produces research-grade graphics. R Markdown → PDF/HTML reports directly. Factor data type for categorical variables.

Why Python over R for ML?

Better production deployment (FastAPI, Docker). Richer deep learning ecosystem (PyTorch, TensorFlow). Better for large-scale ETL pipelines. Stronger OS integration and scripting. Larger ML community and job market.

R · Basics — Syntax & Key Differences from Python
# Assignment — use <- (convention) or = (works too)
name  <- "Priya Nair"
marks <- 88.5
batch  = "PGDBDA-2026"      # = also works

# Print
print(marks)                   # [1] 88.5  ← [1] means "first element"
cat("Student:", name, "\n")   # no [1] prefix — like Python print

# Arithmetic — KEY DIFFERENCES from Python!
5 + 3           # 8
10 / 3          # 3.333...  (always true division)
2 ^ 10          # 1024      (^ not ** like Python)
17 %% 5         # 2         (modulo — same as Python)
17 %/% 5        # 3         (integer division)
sqrt(144)       # 12
abs(-7)         # 7
log(100, 10)    # 2  (log base 10)
log(100)        # 4.60...  (natural log — default in R!)
exp(1)          # 2.71828... (e)
ceiling(3.2)    # 4  (round up)
floor(3.8)      # 3  (round down)
round(3.5672, 2)# 3.57

# Type checking
class(marks)       # "numeric"  (R's float)
class(name)        # "character" (R's string)
is.numeric(marks)  # TRUE
is.character(name) # TRUE

# R indexing starts at 1 (NOT 0 like Python!)
v <- c(10, 20, 30, 40, 50)
v[1]    # 10  (first element — index 1!)
v[5]    # 50  (last element)
v[2:4]  # 20 30 40  (INCLUSIVE on both ends — unlike Python!)

Session 16 · R Data Objects & Packages

Vectors, Lists, Matrices, Data Frames — R's elegantly statistical data structures

2h Theory2h Lab2h Self-Learning
🎯 The Core Analogy

R's data structures map directly to statistical thinking. A vector is one column of data (a single variable in SPSS). A data frame is the complete dataset — rows are observations, columns are variables. A list is a filing cabinet that can hold anything: text, numbers, plots, other lists, or even functions.

matrix / array (2D, same type) vector (1D, homogeneous) c(88, 92, 76, 95, 81) all elements must be same type data.frame (2D, mixed types) numeric col 88, 74, 95 marks, cgpa factor col "A","B","A" grade, cut, gender
R · Lab Exercises — All 5 Tasks
# VECTOR — R's most fundamental object
marks <- c(88, 92, 76, 95, 81)    # c() = combine/concatenate

# Lab Q1 — sum, mean, product
sum(marks)        # 432
mean(marks)       # 86.4
prod(marks)       # product of all elements

# Lab Q2 — airquality built-in dataset
data(airquality)
head(airquality)
summary(airquality)
colSums(is.na(airquality))   # count NAs per column

# Lab Q3 — list of dataframes
df_list <- list(
    batch1 = data.frame(name=c("Priya","Raj"),  marks=c(88,74)),
    batch2 = data.frame(name=c("Zara","Arun"), marks=c(95,81))
)
# Lab Q4 — access each dataframe from list
print(df_list[["batch1"]])    # by name
print(df_list[[2]])          # by index

# Lab Q5 — create, summarise, sort, add column, export
students <- data.frame(
    name   = c("Priya", "Raj", "Zara", "Arun"),
    marks  = c(88, 74, 95, 81),
    passed = c(TRUE, TRUE, TRUE, TRUE)
)
summary(students)                              # statistical summary
str(students)                                  # structure and data types
students$grade <- ifelse(students$marks>=85, "A", "B")
students_sorted <- students[order(-students$marks), ]

# Export to Excel
# install.packages("writexl")
library(writexl)
write_xlsx(students_sorted, "students_report.xlsx")

Session 17 · tidyverse & Data Manipulation

dplyr · ggplot2 · tidyr · rvest — the modern R ecosystem for data science

2h Theory2h Lab3h Self-Learning
🎯 The Core Analogy

tidyverse is a LEGO system for data. Each package (dplyr, tidyr, ggplot2) is a compatible brick that snaps together. The pipe %>% connects them: "Take this data, THEN filter, THEN group, THEN summarise, THEN plot." You read it left to right like English — no nested parentheses, no intermediate variables to track.

R · dplyr + ggplot2 + rvest — All Lab Tasks
library(tidyverse)   # loads dplyr, ggplot2, tidyr, readr, etc.

═══ The 5 Core dplyr Verbs ══════════════════════════════════
diamonds %>%
    filter(cut == "Premium", price > 5000) %>% # 1. filter rows
    select(carat, cut, color, price) %>%          # 2. pick columns
    mutate(ppc = price / carat) %>%               # 3. add/transform col
    arrange(desc(ppc)) %>%                        # 4. sort
    head(10)                                       # top 10 rows

# group_by + summarise = SQL GROUP BY + aggregate
diamonds %>%
    group_by(cut) %>%
    summarise(
        count     = n(),
        avg_price = mean(price),
        max_p     = max(price)
    ) %>%
    arrange(desc(avg_price))

═══ Lab Q3: Pie Chart + Bar Chart ═══════════════════════════
# Bar chart
ggplot(diamonds, aes(x=cut, fill=cut)) +
    geom_bar() +
    labs(title="Diamond Cut Distribution", x="Cut", y="Count") +
    theme_minimal() +
    theme(legend.position="none")

# Pie chart (bar + coord_polar trick)
cut_counts <- diamonds %>% count(cut)
ggplot(cut_counts, aes(x="", y=n, fill=cut)) +
    geom_col(width=1, color="white") +
    coord_polar("y", start=0) +    # this transforms bar → pie
    labs(title="Diamond Cut Proportions") +
    theme_void()

═══ Lab Q1: Load JSON and XML ════════════════════════════════
library(jsonlite); library(XML)
json_df <- fromJSON("data.json", flatten=TRUE)
xml_df  <- xmlToDataFrame(xmlParse("data.xml"))
summary(json_df); summary(xml_df)

═══ Lab Q2: Web Scraping with rvest ═════════════════════════
library(rvest)
page   <- read_html("https://books.toscrape.com")
titles <- page %>% html_nodes("h3 a") %>% html_attr("title")
prices <- page %>% html_nodes(".price_color") %>% html_text()
books  <- data.frame(title=titles, price=prices)
head(books, 5)

Session 18 · R Functions & R Markdown

Writing reusable functions · ChickWeight case study · Reproducible reports

2h Theory2h Lab2h Self-Learning
🎯 R Markdown Analogy

R Markdown is a lab notebook that runs its own experiments. You write prose, embed R code chunks, and when you click "Knit" in RStudio, every chunk executes and its output — tables, charts, statistics — appears right there in the final PDF or HTML. No copy-pasting from console. The report is the reproducible code.

R · Functions + ChickWeight Case Study — Complete Lab Q1
# User-defined function
calculate_gpa <- function(m1, m2, m3) {
    gpa <- (1/3)*m1 + (1/2)*m2 + (1/4)*m3
    cat(sprintf("GPA = %.2f\n", gpa))
    return(gpa)
}
calculate_gpa(80, 75, 90)   # GPA = 79.17

═══ ChickWeight Case Study ═══════════════════════════════════
data(ChickWeight)
str(ChickWeight)    # weight, Time, Chick, Diet

# (a) Weight vs Time for Chick 34
chick34 <- ChickWeight[ChickWeight$Chick == 34, ]
plot(chick34$Time, chick34$weight,
     type  = "b",          # b = both lines and points
     col   = "#b34a22",
     pch   = 16,           # filled circles
     xlab  = "Days after birth",
     ylab  = "Weight (grams)",
     main  = "Chick 34 — Growth Over Time")

# (b) Boxplot for Diet group 4
diet4 <- subset(ChickWeight, Diet == 4)
boxplot(weight ~ Time, data = diet4,
        col = "lightblue",
        main = "Diet 4: Weight by Time Point")

# (c) Mean weight per time for Diet 4
mean_d4 <- tapply(diet4$weight, diet4$Time, mean)
plot(names(mean_d4), mean_d4, type="b",
     col="#b34a22", pch=16,
     xlab="Time", ylab="Mean Weight (g)",
     main="Mean Weight — Diet 4 vs 2")

# (d) Add Diet 2 line to same plot
diet2   <- subset(ChickWeight, Diet == 2)
mean_d2 <- tapply(diet2$weight, diet2$Time, mean)
lines(names(mean_d2), mean_d2, col="#1a6e5c", type="b", pch=17)

# (e) Legend and title
legend("topleft",
       legend = c("Diet 4", "Diet 2"),
       col    = c("#b34a22", "#1a6e5c"),
       lty = 1, pch = c(16,17), bty="n")
R Markdown — Structure of a Report
R Markdown (.Rmd) — Knit to HTML with Ctrl+Shift+K
---
title:   "PG-DBDA Data Analysis Report"
author:  "Your Name — C-DAC Kharghar"
date:    "`r Sys.Date()`"
output:
  html_document:
    toc: true
    toc_float: true
    theme: flatly
---

## Introduction

Analyse the **ChickWeight** dataset to compare dietary groups.

```{r setup, include=FALSE}
knitr::opts_chunk$set(echo=TRUE, warning=FALSE, message=FALSE)
library(tidyverse)
```

```{r growth-chart, fig.width=9, fig.height=5}
data(ChickWeight)
ggplot(ChickWeight, aes(x=Time, y=weight,
       color=factor(Diet), group=Chick)) +
  geom_line(alpha=0.4) +
  stat_summary(aes(group=Diet), fun=mean,
               geom="line", linewidth=1.5) +
  labs(title="Chick Growth by Diet Group", color="Diet") +
  theme_minimal()
```

## Conclusion

Diet 4 shows the highest mean weight gain by day 21 (p < 0.05).

# Inline R: "The dataset has `r nrow(ChickWeight)` observations."
🧪 Which single R package loads dplyr, ggplot2, tidyr, readr, and more together?