Python & R Programming — Complete Revision Guide
C-DAC Kharghar · PG-DBDA February 2026 · Beginner → Advanced
This guide covers every session of the PG-DBDA Python & R module at C-DAC Kharghar. Each topic is explained with real-world Indian analogies, illustrated with SVG diagrams, and backed by working, exam-ready code. Whether you are encountering programming for the first time or brushing up before the lab exam, read each session from the top — concepts build on each other deliberately.
| Component | Weightage | What it tests |
|---|---|---|
| Theory Exam | 40% | Concepts, definitions, tracing code, explaining data structures |
| Lab Exam | 40% | Writing programs, debugging, output prediction |
| Internal | 20% | Assignments, quizzes, mini-projects, class participation |
Sessions 1–2 · Python Foundations
Variables · Data Types · if/else · for/while loops · Strings · Input/Output
Sessions 3–4 · Core Data Structures
Dictionaries · Lists · Comprehensions · Functional tools (map, filter, zip)
Sessions 5–7 · Functions & Tuples
Lambda · Recursion · Modules · Pickling · Tuples as fixed records
Sessions 8–10 · Advanced Python
OOP · Generators · Decorators · Regex · Exception Handling · Logging
Sessions 11–13 · Data Science Toolkit
NumPy · Pandas · Matplotlib · Seaborn · Plotly · BeautifulSoup · Pillow
Session 14 · Python + Databases
SQLite · MySQL · Pandas + SQL · CRUD operations · Parameterised queries
Sessions 15–18 · R Programming
Basics · Vectors · Data Frames · tidyverse · ggplot2 · R Markdown
🔀 Control Flow
if/else, for, while, break, continue — decision and repetition
Session 1📝 Strings
Slicing, methods, palindromes, pangrams, f-strings
Session 2📚 Dictionaries
Key-value, comprehension, ROT-13 cipher lab
Session 3⚙️ Functions
Lambda, map, recursion, *args, closures, pickling
Sessions 5–6🏗️ OOP
Classes, inheritance, generators, decorators, regex
Sessions 8–9🐼 Pandas & NumPy
DataFrames, arrays, groupby, diamonds dataset
Session 11📊 R Basics
Vectors, assignment, arithmetic, RStudio exploration
Session 15🌊 tidyverse
dplyr pipes, ggplot2 charts, data reshaping
Session 17Session 1 · Python Basics & Control Flow
Variables · Data Types · Operators · if/else · for · while · break/continue/pass
Python is like giving instructions to a very literal-minded cook. You write a recipe (program), the cook follows it exactly — no assumptions, no shortcuts. An if-else is: "If the milk is sour, skip it; otherwise use it." A for loop is: "For each student in the class roll call, say their name." The cook never improvises — that precision is Python's strength.
How Python runs your code
Python is an interpreted language — unlike C or Java which compile to machine code first, Python reads your file and executes it line by line. Understanding this journey helps you understand import errors, bytecode caching, and performance characteristics.
| Type | Example | C-DAC Real-world Usage | Type Check |
|---|---|---|---|
int | roll = 2601 | Roll number, batch count, port number | type(roll) |
float | cgpa = 8.75 | CGPA, salary (₹), ML model accuracy | isinstance(x, float) |
str | name = "Priya" | Names, SQL queries, file paths, JSON keys | type(x) == str |
bool | passed = True | Pass/Fail flags, ON/OFF switches, DB flags | isinstance(x, bool) |
NoneType | result = None | Missing data, NULL from DB, unset variable | result is None |
complex | z = 3+4j | Signal processing, DSP, FFT | z.real, z.imag |
# Python is DYNAMICALLY TYPED — no type declaration needed name = "Priya Nair" # str roll = 2601 # int cgpa = 8.92 # float is_placed = True # bool dept = None # NoneType — not yet assigned # Python infers the type at runtime print(type(cgpa)) # <class 'float'> # Multiple assignment on one line x, y, z = 10, 20, 30 # Swap without a temp variable (very Pythonic!) x, y = y, x # x=20, y=10 — tuple packing/unpacking # Type conversion (casting) int("42") # → 42 (string to int) str(3.14) # → "3.14" (float to string) float(7) # → 7.0 (int to float) bool(0) # → False (0, "", [], None are all Falsy) bool("CDAC") # → True (any non-empty string is Truthy)
Control flow tells Python which path to take through your code. Python uses indentation (4 spaces) to define code blocks — no curly braces like Java or C. Wrong indentation is always a SyntaxError.
marks = int(input("Enter marks (0-100): ")) if marks >= 90: print("O — Outstanding") elif marks >= 75: print("A — Distinction") elif marks >= 60: print("B — First Class") elif marks >= 40: print("C — Pass") else: print("F — Fail. Please reappear.") # One-liner ternary — for simple True/False choices result = "Pass" if marks >= 40 else "Fail" print(f"Result: {result}")
# Q1: for loop — print 0! to 10! f, n = 1, 0 for _ in range(11): print(f"{n}! = {f}") n += 1 f *= n # Output: 0!=1, 1!=1, 2!=2, ... 10!=3,628,800 # Q2: while loop — all factorials under 2 billion f, n = 1, 0 while f < 2_000_000_000: # underscore separators for readability print(f"{n}! = {f:,}") # :, adds comma formatting to numbers n += 1 f *= n
| Keyword | Effect | Analogy | Common Use |
|---|---|---|---|
break | Exit entire loop now | Emergency stop on a machine | Found what you searched for |
continue | Skip this iteration, go to next | Skip a corrupted track on Spotify | Skip invalid data rows |
pass | Do nothing — placeholder | Empty room stub in a blueprint | Skeleton class/function |
print(type(True)) output?Session 2 · Strings
Text manipulation — the single most critical skill for data cleaning and preprocessing
A string is a Mumbai local train with coaches numbered from 0. Coach 0 is the first, coach -1 is the last (counting from the back). You can board at any coach (s[3]), or take a stretch of coaches (s[2:7]), or walk the entire train backwards (s[::-1]). The train is immutable — you can not rebuild a coach in place. s[0] = 'X' gives a TypeError; you must produce a new string.
s = "C-DAC Kharghar Navi Mumbai" # SLICING — [start:stop:step] (stop is exclusive!) print(s[0:5]) # 'C-DAC' print(s[6:]) # 'Kharghar Navi Mumbai' print(s[::-1]) # reverse the whole string print(s[6:14:2]) # every 2nd char from 6 to 14 # IMPORTANT METHODS print(s.upper()) # 'C-DAC KHARGHAR NAVI MUMBAI' print(s.lower()) # 'c-dac kharghar navi mumbai' print(s.replace('Navi','New'))# replace substring print(s.split(' ')) # ['C-DAC', 'Kharghar', 'Navi', 'Mumbai'] print(', '.join(['a','b','c'])) # 'a, b, c' print(s.strip()) # remove leading/trailing whitespace print(s.count('a')) # count occurrences of 'a' print(s.find('Kharghar')) # index position or -1 if not found print(s.startswith('C-DAC'))# True print(s.endswith('Mumbai')) # True # F-STRINGS — modern, fast, readable string formatting batch, cgpa = 2026, 8.9247 print(f"PG-DBDA {batch} — C-DAC Kharghar") print(f"CGPA: {cgpa:.2f}") # 'CGPA: 8.92' (2 decimal places) print(f"Rank: {1:04d}") # 'Rank: 0001' (zero-padded) print(f"Total: ₹{75000:,}") # 'Total: ₹75,000'
import re # Q3: Phrase palindrome checker def is_palindrome(phrase): """Strip non-letters, lowercase, compare to reverse.""" clean = re.sub(r'[^a-zA-Z]', '', phrase).lower() return clean == clean[::-1] tests = ["Was it a rat I saw?", "Step on no pets", "Dammit, I'm mad!", "Hello CDAC"] for t in tests: print(f"{'✓' if is_palindrome(t) else '✗'} {t!r}") # Q4: Pangram checker (contains all 26 English letters) def is_pangram(sentence): letters = set(sentence.lower()) return set('abcdefghijklmnopqrstuvwxyz').issubset(letters) print(is_pangram("The quick brown fox jumps over the lazy dog")) # True print(is_pangram("Hello World")) # False
Session 3 · Dictionaries
Key-Value stores — the data structure powering Big Data, JSON, and every web API
A Python dict is exactly like your phone's contact book. You search by name (key) and get the phone number (value) instantly — you do not scroll through every contact. Lookup is O(1) whether you have 10 or 10 million entries. This is the same principle behind Hadoop's shuffle-sort phase, Redis, and MongoDB document stores.
# CREATE student = { "name": "Priya Nair", "roll": 2601, "marks": {"Python": 88, "R": 79}, # nested dict! "batch": "PGDBDA-Feb2026" } # ACCESS — two ways print(student["name"]) # KeyError if key missing print(student.get("dept", "N/A")) # safe — returns default # MODIFY & ADD student["marks"]["Java"] = 83 # add to nested dict student["graduated"] = False # add new top-level key # DELETE del student["roll"] popped = student.pop("batch") # remove and return value # ITERATION for key, val in student.items(): print(f" {key}: {val}") # DICT COMPREHENSION — like list comp but for dicts squares = {x: x**2 for x in range(1, 6)} # {1:1, 2:4, 3:9, 4:16, 5:25} # ── ROT-13 CIPHER LAB Q1 ────────────────────────────────────── def rot13(text): upper = 'ABCDEFGHIJKLMNOPQRSTUVWXYZ' lower = upper.lower() table = str.maketrans( upper + lower, upper[13:] + upper[:13] + lower[13:] + lower[:13] ) return text.translate(table) secret = "Pnrfne pvcure? V zhpu cersre Pnrfne fnynq!" print(rot13(secret)) # Caesar cipher? I much prefer Caesar salad! print(rot13(rot13(secret)) == secret) # True — self-inverse!
Session 4 · Lists
Python's most versatile data structure — ordered, mutable, and infinitely flexible
A list is like a Zomato order cart. You can add items (append), remove one (remove), look at what's at the top ([-1]), merge two carts (+), and even create a mini-cart of just the first three items ([:3]). Everything is ordered by position, and nothing is fixed — it is completely mutable.
scores = [88, 72, 95, 61, 84, 91] --- ADDING --- scores.append(78) # add to end scores.insert(0, 100) # insert at index 0 scores.extend([55, 67]) # merge another list scores += [82] # shorthand extend --- REMOVING --- scores.remove(61) # remove first occurrence of value 61 scores.pop() # remove & return last element scores.pop(2) # remove & return element at index 2 del scores[0:2] # delete a slice scores.clear() # empty the entire list --- SEARCHING & STATS --- print(max(scores), min(scores), sum(scores)) print(scores.index(95)) # position of 95 print(scores.count(88)) # how many times 88 appears print(95 in scores) # True — membership test --- SORTING --- scores.sort() # ascending, modifies IN-PLACE scores.sort(reverse=True) # descending, in-place ranked = sorted(scores) # returns NEW sorted list, original unchanged --- LIST COMPREHENSION (Pythonic & fast!) --- passed = [s for s in scores if s >= 40] doubled = [s * 2 for s in scores] grades = ['P' if s>=40 else 'F' for s in scores] --- CRITICAL: Copy vs Reference --- a = [1, 2, 3] b = a # SAME list! b IS a — dangerous! c = a.copy() # true independent copy d = a[:] # slice copy — also independent e = list(a) # another way to copy # Test: b.append(99) → a also has 99. c.append(99) → a unchanged.
Most common beginner mistake: b = a does NOT copy a list — it makes b point to the same list in memory as a. If you do b.append(99), you will find 99 in a as well. Always use a.copy() or a[:] when you need an independent copy.
Sessions 5–6 · Functions & Modules
Lambda · Recursion · *args/**kwargs · Modules · Pickling — reusable code building blocks
A function is a mixer-grinder. You put ingredients in (parameters), it processes, you get output (return value). Once you own it, use it anytime (reusability). A lambda is a small hand blender — quick, one-time use, no need to give it a shelf. A module is the entire appliance catalogue from a brand — grouped, importable, ready to use.
# 1. Named function with default argument def greet(name, batch="PGDBDA-2026"): return f"Welcome {name} to C-DAC Kharghar — {batch}!" greet("Priya") # uses default batch greet("Raj", "AI-ML-2026") # overrides default # 2. *args — variable positional arguments (packed as tuple) def total_marks(*subjects): print(f"Got {len(subjects)} subjects") return sum(subjects) total_marks(88, 79, 91, 83) # 341 # 3. **kwargs — variable keyword arguments (packed as dict) def student_info(**details): for k, v in details.items(): print(f" {k}: {v}") student_info(name="Zara", roll=2610, cgpa=9.1) # 4. Lambda — anonymous, one-liner, no return keyword square = lambda x: x**2 is_even = lambda n: n % 2 == 0 # Lambda with sorted() — sort students by marks (highest first) students = [("Priya",88), ("Raj",74), ("Zara",95), ("Dev",81)] ranked = sorted(students, key=lambda s: s[1], reverse=True) # [('Zara', 95), ('Priya', 88), ('Dev', 81), ('Raj', 74)] # 5. Recursive function (Lab Q2 — factorial) def factorial(n): if n == 0: return 1 # BASE CASE — must exist! return n * factorial(n - 1) # recursive case for i in range(1, 11): print(f"{i:2d} {factorial(i):>10,}")
import pickle model_data = { "batch": "PGDBDA-2026", "students": ["Priya", "Raj", "Zara"], "accuracy": 0.94 } # Serialize (pack) to disk — wb = write binary with open("batch.pkl", "wb") as f: pickle.dump(model_data, f) # Deserialize (unpack) — rb = read binary with open("batch.pkl", "rb") as f: loaded = pickle.load(f) print(loaded["students"]) # ['Priya', 'Raj', 'Zara'] # Used heavily in ML to save/load trained scikit-learn models!
Session 7 · Tuples — Deep Dive
Immutable sequences — faster, safer, hashable. The fixed-schema record of Python.
A tuple is a sealed Aadhaar card. Once issued, the data is fixed — you cannot change your date of birth in it. A list is an editable form. Use tuples for fixed facts: GPS coordinates (18.92, 72.83), student record (2601, "Priya", 88.5), RGB colour (255, 0, 0), DB row. These facts should never change in your program.
Use Tuple when…
Data is fixed and should never change: DB rows, coordinates, RGB values, function returning multiple values, dictionary keys (tuples are hashable — lists are not)
Use List when…
Data grows or shrinks: exam scores accumulating, a shopping cart, log entries, any collection that changes during the program's lifetime
# Create coords = (18.9220, 72.8347) # lat/lon near C-DAC Kharghar area rgb = (255, 140, 0) # dark orange single = (42,) # MUST have trailing comma for single-element! t = (10, 20, 30, 20, 10) # Lab Q1: Find repeated items repeated = [x for x in set(t) if t.count(x) > 1] print(repeated) # [10, 20] # Lab Q2: Sort list of tuples by float element items = [('item1', '12.20'), ('item2', '15.10'), ('item3', '24.5')] items_sorted = sorted(items, key=lambda x: float(x[1]), reverse=True) # [('item3','24.5'), ('item2','15.10'), ('item1','12.20')] # Lab Q3: Count elements until a tuple is found mixed = [10, 20, 30, (40, 50), 60] count = next(i for i, x in enumerate(mixed) if isinstance(x, tuple)) print(count) # 3 # Lab Q4: Element-wise sum using zip() t1, t2, t3 = (1,2,3,4), (3,5,2,1), (2,2,3,1) result = tuple(a+b+c for a,b,c in zip(t1, t2, t3)) print(result) # (6, 9, 8, 6) # Tuple unpacking — very Pythonic lat, lon = coords x, y, *rest = (1, 2, 3, 4, 5) # rest = [3,4,5]
Sessions 8–9 · OOP, Generators, Decorators & Regex
The professional Python toolkit — how every production system is actually built
A Class is the blueprint; an Object is the actual building. C-DAC has one blueprint called Student with attributes (name, roll, marks) and methods (calculate_gpa, display). Every enrolled student — Priya, Raj, Zara — is an object created from that blueprint. Same blueprint, thousands of instances, each with their own data.
class Student: course = "PG-DBDA" # class variable — shared by ALL objects def __init__(self, roll, name, marks_dict): # Instance variables — unique to EACH object self.roll = roll self.name = name self.marks = marks_dict # {subject: [list of 5 marks]} def __str__(self): # called by print(student) return f"[{self.roll}] {self.name} — {self.course}" def calculate_gpa(self, m1, m2, m3): # GPA formula from syllabus return (1/3)*m1 + (1/2)*m2 + (1/4)*m3 def has_failed(self): # any() short-circuits — stops at first True found return any( m < 40 for mlist in self.marks.values() for m in mlist ) # Inheritance — Faculty extends Person class Person: def __init__(self, name): self.name = name def speak(self): return "..." class Faculty(Person): # inherits Person def speak(self): # method OVERRIDING return f"Prof. {self.name}" # Polymorphism — same method call, different behaviour people = [Faculty("Dr. Mehta"), Person("Visitor")] for p in people: print(p.speak())
Why Generators matter for Big Data: A list of 10 million rows takes ~400MB RAM. A generator produces one item at a time, using barely any memory — which is exactly how Spark processes huge datasets in chunks. yield pauses the function and returns one value; the next call resumes from where it stopped.
# Generator — memory-efficient sequence class NumberSeries: def even_numbers(self, limit): n = 0 while n <= limit: yield n # pauses here; next() call resumes n += 2 ns = NumberSeries() for num in ns.even_numbers(20): print(num, end=" ") # 0 2 4 6 8 10 12 14 16 18 20 # Decorator — wraps a function to add behaviour WITHOUT changing it # Analogy: a SmartWatch wraps your wrist — your wrist is unchanged import functools, time def log_method_call(func): @functools.wraps(func) # preserves original function name def wrapper(*args, **kwargs): print(f"→ Calling {func.__name__}({args[1:]})") start = time.time() result = func(*args, **kwargs) print(f"← Done in {time.time()-start:.4f}s") return result return wrapper class Student: @log_method_call # apply decorator with @ def calculate_gpa(self, m1, m2, m3): return (1/3)*m1+(1/2)*m2+(1/4)*m3 s = Student() s.calculate_gpa(80, 75, 90) # → Calling calculate_gpa((80, 75, 90)) # ← Done in 0.0001s
Session 10 · Exception Handling & Logging
Writing code that survives the real world — graceful failure over catastrophic crashes
Exception handling is an emergency protocol at a hospital. Things you don't expect will go wrong. You don't prevent every scenario — you have trained responders (except clauses), cleanup teams (finally blocks), and post-incident reports (logging). Without this, one unexpected user input can bring down an entire production system running FRAS or Digi-Exam.
try: marks = int(input("Enter marks: ")) # ValueError if "abc" avg = 500 / marks # ZeroDivisionError if 0 print(f"Average: {avg:.1f}") except ValueError: print("Please enter a valid integer.") except ZeroDivisionError: print("Marks cannot be zero.") except Exception as e: print(f"Unexpected: {type(e).__name__}: {e}") else: print("Calculation successful.") # runs ONLY if no exception finally: print("Done.") # ALWAYS runs (cleanup) # CUSTOM EXCEPTION — for domain-specific validation class InvalidMarksError(Exception): def __init__(self, val): super().__init__(f"Marks {val} out of range (0–100)") def validate_marks(m): if not 0 <= m <= 100: raise InvalidMarksError(m) # explicitly raise return m # LOGGING — structured, persistent, production-grade import logging logging.basicConfig( level = logging.DEBUG, format = '%(asctime)s [%(levelname)s] %(message)s', handlers = [ logging.FileHandler('cdac_app.log'), logging.StreamHandler() # also print to console ] ) logging.info("FRAS v2.0 started — C-DAC Kharghar") logging.warning("DB connection latency > 500ms") logging.error("Face recognition model load failed") logging.critical("Disk full — attendance records at risk!")
| Level | Number | When to use | Example |
|---|---|---|---|
DEBUG | 10 | Detailed tracing during development | "Processing student ID 2601" |
INFO | 20 | Normal operation milestones | "API server started on port 8000" |
WARNING | 30 | Unexpected but recoverable | "Disk usage at 85%" |
ERROR | 40 | Serious failure, function aborted | "Cannot connect to PostgreSQL" |
CRITICAL | 50 | System about to die | "Out of memory — killing process" |
Session 11 · Pandas & NumPy
The data science powerhouses — arrays, DataFrames, wrangling, and the diamonds dataset lab
NumPy is a scientific calculator; Pandas is Excel inside Python. NumPy works on entire arrays at once — no Python loops, C-speed math. Pandas gives you labelled rows and columns (like an Excel sheet) with SQL-like groupby, merge, and pivot. Together they turn weeks of manual spreadsheet work into a few lines of code.
import numpy as np import pandas as pd ═══ NumPy Essentials ═══════════════════════════════════════ marks = np.array([88, 92, 76, 95, 61, 84]) # Vectorised — all elements at once print(marks + 5) # add 5 to every element print(marks[marks > 85]) # [88, 92, 95] — boolean indexing print(marks.mean(), marks.std()) # statistics print(np.percentile(marks, [25,50,75])) # Image array (Lab Q2) from PIL import Image img = Image.open('photo.jpg') arr = np.array(img) # (H, W, 3) array restored = Image.fromarray(arr) ═══ Pandas — Diamonds Dataset (Lab Q4) ═════════════════════ df = pd.read_csv('diamonds.csv') # Print first 6 rows print(df.head(6)) # Mean of each numeric column print(df.mean(numeric_only=True)) # Count / min / max price for each cut print(df.groupby('cut')['price'].agg(['count','min','max'])) # Concise summary df.info() # dtypes, non-null counts df.describe() # mean, std, quartiles # Count duplicate rows print(df.duplicated().sum()) # Data cleaning df.dropna(inplace=True) # drop rows with NaN df.drop_duplicates(inplace=True) # Add computed column df['price_per_carat'] = df['price'] / df['carat'] # Filter premium = df[(df['cut']=='Ideal') & (df['price']>5000)]
Session 12 · Visualisation & Web Scraping
matplotlib · seaborn · plotly · ggplot · BeautifulSoup4 — making data speak visually
Visualisation is the translation layer between numbers and human understanding. A table of 10,000 rows tells you nothing at a glance. A heatmap of correlations tells you everything in 3 seconds. The best data scientists spend more time on charts than on models — because a wrong model caught by a good plot saves weeks of wasted computation. BeautifulSoup is how you collect that data from the web in the first place.
| Library | Best Charts | Strength | Weakness |
|---|---|---|---|
matplotlib | Any — line, bar, scatter, hist, pie, 3D | Absolute control, publication-ready, subplots | Verbose code for simple charts |
seaborn | heatmap, boxplot, violin, pairplot, distplot | Statistical defaults, beautiful out-of-box | Less flexible than matplotlib |
plotly | Any + 3D, choropleth, sunburst, candlestick | Interactive (hover, zoom, pan), web embed | Heavy file size, slower render |
plotnine | R-style ggplot grammar in Python | Layered aesthetic mapping, familiar to R users | Slower, smaller community |
df.plot() | Basic line, bar, hist, scatter, box | Zero extra import — just call on DataFrame | Limited customisation |
Part 1 — matplotlib: Anatomy & Core Plots
matplotlib works on a Figure → Axes → Plot hierarchy. The Figure is the entire canvas. The Axes is one chart area inside it (you can have multiple). Every element — title, labels, ticks, gridlines, legend — is individually controllable.
import matplotlib.pyplot as plt import numpy as np ═══ 1. Line Chart ═══════════════════════════════════════════ months = ['Jan','Feb','Mar','Apr','May','Jun'] sales = [120, 145, 132, 178, 190, 210] fig, ax = plt.subplots(figsize=(9, 4)) ax.plot(months, sales, marker='o', color='#b34a22', linewidth=2, label='Monthly Sales') ax.fill_between(months, sales, alpha=0.1, color='#b34a22') ax.set_title('C-DAC Sales Dashboard') ax.set_xlabel('Month'); ax.set_ylabel('Units Sold') ax.legend(); ax.grid(True, alpha=0.3) plt.tight_layout(); plt.show() ═══ 2. Bar Chart ════════════════════════════════════════════ subjects = ['Python', 'R', 'Stats', 'Java', 'Linux'] avg_marks = [82, 76, 88, 71, 85] colors = ['#b34a22' if m>=80 else '#d4c8b0' for m in avg_marks] fig, ax = plt.subplots(figsize=(8, 4)) bars = ax.bar(subjects, avg_marks, color=colors, edgecolor='white') ax.bar_label(bars, fmt='%d') # add value labels on bars ax.axhline(80, color='#1a6e5c', linestyle='--', label='Pass Line') ax.set_title('PG-DBDA Batch Average Marks') ax.set_ylim(0, 100); ax.legend() plt.tight_layout(); plt.show() ═══ 3. Histogram ════════════════════════════════════════════ marks_data = np.random.normal(72, 12, 200) # simulate marks distribution fig, ax = plt.subplots(figsize=(8, 4)) ax.hist(marks_data, bins=20, color='#1a6e5c', edgecolor='white', alpha=0.8) ax.axvline(marks_data.mean(), color='#b34a22', linestyle='--', label=f'Mean={marks_data.mean():.1f}') ax.set_title('Marks Distribution'); ax.legend() plt.tight_layout(); plt.show() ═══ 4. Subplot Grid — show multiple charts at once ══════════ fig, axes = plt.subplots(1, 3, figsize=(14, 4)) axes[0].plot([1,2,3], [4,5,6]); axes[0].set_title('Line') axes[1].bar(['A','B','C'], [3,7,5]); axes[1].set_title('Bar') axes[2].scatter([1,2,3], [3,1,4]); axes[2].set_title('Scatter') plt.suptitle('Multi-panel Dashboard') plt.tight_layout(); plt.show() ═══ 5. Save figure to file ══════════════════════════════════ fig.savefig('report_chart.png', dpi=150, bbox_inches='tight') fig.savefig('report_chart.pdf') # vector format for reports
Part 2 — seaborn: Statistical Visualisation
seaborn is built on top of matplotlib but adds a higher-level interface designed specifically for statistical graphics. It understands Pandas DataFrames natively — pass column names as strings and seaborn does the grouping, colouring, and labelling for you.
import seaborn as sns import matplotlib.pyplot as plt import pandas as pd import numpy as np ═══ Seaborn Themes — set once, applies everywhere ═══════════ sns.set_theme(style='whitegrid', palette='muted') # Other styles: darkgrid, white, ticks, dark ═══ 1. Heatmap — Lab Q2 Step 1 & 2 ════════════════════════ df = pd.read_csv('data.csv') corr = df.corr(numeric_only=True) plt.figure(figsize=(10, 8)) sns.heatmap( corr, annot = True, # show values in cells fmt = '.2f', # 2 decimal places cmap = 'RdYlGn', # Red-Yellow-Green diverging center = 0, # 0 = white centre linewidths= 0.5, # cell borders square = True, # force square cells vmin=-1, vmax=1 # fix scale -1 to +1 ) plt.title('Correlation Matrix — PG-DBDA Dataset', pad=14) plt.tight_layout(); plt.show() ═══ 2. Find highest correlated pair (Lab Q2 Step 3) ════════ mask = np.triu(np.ones_like(corr, dtype=bool)) # upper triangle col1, col2 = corr.mask(mask).stack().idxmax() print(f"Highest correlation: {col1} ↔ {col2}") ═══ 3. Scatter plot (Lab Q2 Step 4) ════════════════════════ fig, ax = plt.subplots(figsize=(8, 6)) sns.scatterplot(data=df, x=col1, y=col2, alpha=0.6, color='#b34a22', ax=ax) sns.regplot(data=df, x=col1, y=col2, scatter=False, color='#1a6e5c', ax=ax, label='Trend line') ax.set_title(f'{col1} vs {col2} (r={corr.loc[col1,col2]:.2f})') ax.legend() plt.tight_layout(); plt.show() ═══ 4. Box Plot — compare distributions by category ════════ tips = sns.load_dataset('tips') # built-in sample dataset plt.figure(figsize=(9, 5)) sns.boxplot(data=tips, x='day', y='total_bill', hue='sex') plt.title('Bill Distribution by Day and Gender') plt.show() ═══ 5. Violin Plot — like boxplot but shows distribution shape plt.figure(figsize=(9, 5)) sns.violinplot(data=tips, x='day', y='total_bill', hue='sex', split=True, palette='muted') plt.title('Bill Distribution — Violin (shows density shape)') plt.show() ═══ 6. Pair Plot — all columns vs all columns at once ══════ iris = sns.load_dataset('iris') sns.pairplot(iris, hue='species', diag_kind='kde') plt.suptitle('Iris Dataset — Pairwise Relationships', y=1.02) plt.show() ═══ 7. Count Plot — categorical frequency bar ═══════════════ sns.countplot(data=tips, x='day', order=['Thur','Fri','Sat','Sun']) plt.title('Number of Customers per Day') plt.show()
seaborn vs matplotlib — key rule: Start with seaborn for speed. When you need to customise something seaborn can't do, use ax.set_*() methods on the underlying matplotlib Axes object that seaborn returns. They work together seamlessly.
Part 3 — plotly: Interactive Visualisation
plotly produces charts that live in a browser — you can hover, zoom, pan, click to filter, and export as PNG. This is what you embed in Dash dashboards, Flask web apps, and HTML reports. The plotly.express module gives you the same one-liner simplicity as seaborn, but interactive.
import plotly.express as px import plotly.graph_objects as go import pandas as pd ═══ 1. Interactive Scatter with trendline (Lab Q2) ══════════ df = pd.read_csv('data.csv') fig = px.scatter( df, x=col1, y=col2, trendline = "ols", # OLS regression line color = 'category', # colour by category column hover_data = ['id'], # show extra info on hover title = f"{col1} vs {col2} (Interactive)" ) fig.update_layout(template='plotly_white') fig.show() # opens in browser fig.write_html("scatter.html") # save as shareable HTML ═══ 2. Interactive Bar Chart ════════════════════════════════ fig = px.bar( df.groupby('cut')['price'].mean().reset_index(), x='cut', y='price', color='cut', text_auto='.0f', title='Average Diamond Price by Cut' ) fig.show() ═══ 3. Sunburst — hierarchical breakdown ════════════════════ fig = px.sunburst( px.data.tips(), path=['day', 'time', 'sex'], values='total_bill', title='Bill Breakdown (Day → Time → Gender)' ) fig.show() ═══ 4. 3D Scatter ═══════════════════════════════════════════ fig = px.scatter_3d( px.data.iris(), x='sepal_length', y='sepal_width', z='petal_length', color='species', symbol='species', title='Iris Dataset — 3D View' ) fig.show() ═══ 5. Animated Chart — add time dimension ══════════════════ fig = px.scatter( px.data.gapminder(), x='gdpPercap', y='lifeExp', size='pop', color='continent', hover_name='country', animation_frame='year', log_x=True, title='Gapminder — GDP vs Life Expectancy' ) fig.show()
Part 4 — ggplot / plotnine: Grammar of Graphics
The Grammar of Graphics is a principled way to describe charts. Every chart is described as: data + aesthetic mapping + geometric layer + optional scales/facets. This is the native language of R's ggplot2, and Python's plotnine brings this exact grammar to Python. Understanding it makes you fluent in both Python and R visualisation.
# pip install plotnine from plotnine import * import pandas as pd import seaborn as sns diamonds = sns.load_dataset('diamonds') ═══ Basic scatter ════════════════════════════════════════════ p = (ggplot(diamonds, aes(x='carat', y='price', color='cut')) + geom_point(alpha=0.3, size=0.5) + labs(title='Diamond Price vs Carat', x='Carat Weight', y='Price (USD)') + theme_minimal() ) print(p) ═══ Faceted chart — one panel per cut category ═══════════════ p2 = (ggplot(diamonds.sample(2000), aes(x='carat', y='price')) + geom_point(color='#b34a22', alpha=0.5, size=1) + geom_smooth(method='lm', color='#1a6e5c') + facet_wrap('~cut', ncol=3) # one panel per cut + labs(title='Price vs Carat by Cut Quality') + theme_bw() ) print(p2) ═══ Bar + coord_flip ════════════════════════════════════════ p3 = (ggplot(diamonds, aes(x='cut', fill='cut')) + geom_bar() + coord_flip() # horizontal bars + theme_classic() + labs(title='Diamond Cut Counts') ) print(p3)
Part 5 — pandas .plot(): Quick EDA
import pandas as pd import matplotlib.pyplot as plt df = pd.read_csv('diamonds.csv') # Line — index vs value df['price'][:200].plot(title='First 200 Prices'); plt.show() # Histogram of numeric columns df.hist(figsize=(12,8), bins=30, color='#b34a22'); plt.show() # Box plot by group df.boxplot(column='price', by='cut'); plt.show() # Grouped bar chart df.groupby('cut')['price'].mean().plot( kind='bar', color='#1a6e5c', rot=0 ) plt.title('Mean Price by Cut'); plt.show()
Part 6 — BeautifulSoup: Web Scraping
Web scraping is the process of programmatically extracting data from websites. The workflow is always: fetch the HTML → parse it → locate the elements → extract the data → clean and store. BeautifulSoup handles steps 2–4; requests handles step 1; Pandas handles step 5.
import requests from bs4 import BeautifulSoup import pandas as pd import time ═══ Basic Scrape (Lab Q1) ════════════════════════════════════ # Set a header so you look like a real browser headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0)"} resp = requests.get("https://books.toscrape.com", headers=headers) print(resp.status_code) # 200 = success, 403 = blocked, 404 = not found soup = BeautifulSoup(resp.text, 'html.parser') # Find all book titles titles = [t.find('a')['title'] for t in soup.find_all('h3')] prices = [p.text for p in soup.find_all('p', class_='price_color')] ratings= [a['class'][1] for a in soup.find_all('p', class_='star-rating')] for t, p, r in zip(titles[:5], prices[:5], ratings[:5]): print(f"{t[:35]:35s} | {p} | {r} stars") ═══ CSS Selector approach (alternative to find_all) ══════════ # soup.select() uses CSS selectors — very powerful! book_articles = soup.select('article.product_pod') books = [] for article in book_articles: books.append({ 'title' : article.select_one('h3 a')['title'], 'price' : article.select_one('.price_color').text, 'rating': article.select_one('p.star-rating')['class'][1], 'in_stock': 'In stock' in article.select_one('.availability').text }) df_books = pd.DataFrame(books) print(df_books.head()) df_books.to_csv('books_scraped.csv', index=False) ═══ Multi-page Scraping (paginate through all 50 pages) ══════ all_books = [] base_url = "https://books.toscrape.com/catalogue/page-{}.html" for page in range(1, 6): # scrape first 5 pages url = base_url.format(page) resp = requests.get(url, headers=headers) soup = BeautifulSoup(resp.text, 'html.parser') for article in soup.select('article.product_pod'): all_books.append({ 'title': article.select_one('h3 a')['title'], 'price': article.select_one('.price_color').text, 'page': page }) time.sleep(1) # BE POLITE — 1 second delay between pages! print(f"Page {page} done — {len(all_books)} books so far") df_all = pd.DataFrame(all_books) print(f"Total scraped: {len(df_all)} books") ═══ Key Parsing Methods Reference ═══════════════════════════ # soup.find('tag') — first matching tag # soup.find_all('tag') — all matching tags (returns list) # soup.find('tag', class_='x') — tag with specific class # soup.find('tag', id='x') — tag with specific id # soup.select('div.card h2') — CSS selector (most flexible) # soup.select_one('...') — CSS selector, first match only # tag.text / tag.get_text() — extract text content # tag['href'] / tag['class'] — extract attribute value # tag.get('attr', default) — safe attribute access
Web Scraping Ethics & Legality: (1) Always check robots.txt — if the site disallows scraping, do not proceed. (2) Add time.sleep(1) between requests — do not hammer servers. (3) Never scrape personal data (names, emails, phone numbers) without consent. (4) For heavy scraping, use an API if one exists — it is faster, more reliable, and officially permitted. Legal violations can result in banning or litigation.
| Method | What it returns | Example |
|---|---|---|
soup.find('h1') | First <h1> tag object (or None) | soup.find('h1').text |
soup.find_all('a') | List of all <a> tags | [a['href'] for a in soup.find_all('a')] |
soup.select('div.card') | List via CSS selector | soup.select('.product h3 a') |
soup.select_one('p.price') | First match via CSS selector | soup.select_one('.price').text |
tag.text | All text content of tag | p_tag.text.strip() |
tag['attr'] | Attribute value | a_tag['href'], img_tag['src'] |
tag.get('attr', '') | Attribute value with default (safe) | a_tag.get('class', []) |
tag.parent | Parent element | span.parent.find('h2') |
tag.next_sibling | Next sibling element | Navigate DOM tree |
Session 13 · Pillow, Audio & Virtual Environments
Image processing · Audio files · Isolated Python environments for production
═══ Pillow — Image Lab ═════════════════════════════════════ from PIL import Image import numpy as np img = Image.open('photo.jpg') print(img.size, img.mode) # (1920, 1080) RGB # Convert to NumPy array (Lab Q2) arr = np.array(img) print(arr.shape) # (1080, 1920, 3) — H×W×channels print(arr[0, 0]) # [255 255 255] — top-left pixel RGB # Back to image restored = Image.fromarray(arr) restored.save('output.jpg') # Transforms img.resize((800, 600)) # new size img.rotate(90) # rotate 90° img.convert('L') # RGB → Grayscale img.crop((100,100,500,400)) # (left, top, right, bottom) ═══ Virtual Environment — Essential Commands ════════════════ # Create python3 -m venv myproject_env # Activate source myproject_env/bin/activate # Linux / Mac myproject_env\Scripts\activate # Windows PowerShell # Install packages (only inside this env) pip install fastapi uvicorn pandas numpy pillow # Freeze for sharing / deployment pip freeze > requirements.txt # Recreate on another machine / production server pip install -r requirements.txt # Deactivate deactivate
Session 14 · Python + Databases
Connecting Python to SQLite, MySQL, and PostgreSQL — the data pipeline bridge
Python is the waiter; the database is the kitchen. The waiter (Python) takes your order (SQL query), walks it to the kitchen (DB engine), brings the food (result set), and serves it to you. The connection is the bridge. Always close it when done — like clocking out after shift — or you leak resources.
import sqlite3, pandas as pd # Connect (creates file if not exists) conn = sqlite3.connect('cdac_attendance.db') cur = conn.cursor() # Create table cur.execute(""" CREATE TABLE IF NOT EXISTS students ( id INTEGER PRIMARY KEY AUTOINCREMENT, name TEXT NOT NULL, roll TEXT UNIQUE, cgpa REAL DEFAULT 0.0, batch TEXT ) """) # Insert — use ? placeholders (NEVER string format — SQL injection!) batch_data = [ ("Priya Nair", "CDAC2601", 8.9, "PGDBDA-26"), ("Raj Kumar", "CDAC2602", 7.4, "PGDBDA-26"), ("Zara Sheikh", "CDAC2610", 9.1, "PGDBDA-26"), ] cur.executemany( "INSERT OR IGNORE INTO students(name,roll,cgpa,batch) VALUES(?,?,?,?)", batch_data ) conn.commit() # Query directly into Pandas DataFrame — best practice! df = pd.read_sql("SELECT * FROM students WHERE cgpa > 8", conn) print(df) # Always close conn.close() ═══ MySQL Connection ═══════════════════════════════════════ # pip install mysql-connector-python import mysql.connector conn = mysql.connector.connect( host="localhost", user="root", password="cdac@123", database="pgdbda" ) # Identical cursor/execute/commit/close API as SQLite ═══ PostgreSQL Connection ══════════════════════════════════ # pip install psycopg2-binary import psycopg2 conn = psycopg2.connect( host="localhost", port=5432, dbname="fras_db", user="cdac", password="cdac@123" ) # Use %s instead of ? for parameterised queries in psycopg2
f"SELECT * FROM users WHERE name='{user_input}'"?Session 15 · Introduction to R
Statistical computing — built for data by statisticians, beloved by researchers worldwide
If Python is a Swiss Army knife, R is a surgeon's precision scalpel. Python does everything — web, ML, automation, APIs, databases. R is purpose-built for statistical analysis, hypothesis testing, and publication-quality charts. A data scientist at a pharmaceutical company, IIM, or ISRO research lab will almost certainly use R.
Why R over Python for stats?
10,000+ CRAN packages for every statistical method. Built-in distributions and hypothesis tests. ggplot2 produces research-grade graphics. R Markdown → PDF/HTML reports directly. Factor data type for categorical variables.
Why Python over R for ML?
Better production deployment (FastAPI, Docker). Richer deep learning ecosystem (PyTorch, TensorFlow). Better for large-scale ETL pipelines. Stronger OS integration and scripting. Larger ML community and job market.
# Assignment — use <- (convention) or = (works too) name <- "Priya Nair" marks <- 88.5 batch = "PGDBDA-2026" # = also works # Print print(marks) # [1] 88.5 ← [1] means "first element" cat("Student:", name, "\n") # no [1] prefix — like Python print # Arithmetic — KEY DIFFERENCES from Python! 5 + 3 # 8 10 / 3 # 3.333... (always true division) 2 ^ 10 # 1024 (^ not ** like Python) 17 %% 5 # 2 (modulo — same as Python) 17 %/% 5 # 3 (integer division) sqrt(144) # 12 abs(-7) # 7 log(100, 10) # 2 (log base 10) log(100) # 4.60... (natural log — default in R!) exp(1) # 2.71828... (e) ceiling(3.2) # 4 (round up) floor(3.8) # 3 (round down) round(3.5672, 2)# 3.57 # Type checking class(marks) # "numeric" (R's float) class(name) # "character" (R's string) is.numeric(marks) # TRUE is.character(name) # TRUE # R indexing starts at 1 (NOT 0 like Python!) v <- c(10, 20, 30, 40, 50) v[1] # 10 (first element — index 1!) v[5] # 50 (last element) v[2:4] # 20 30 40 (INCLUSIVE on both ends — unlike Python!)
Session 16 · R Data Objects & Packages
Vectors, Lists, Matrices, Data Frames — R's elegantly statistical data structures
R's data structures map directly to statistical thinking. A vector is one column of data (a single variable in SPSS). A data frame is the complete dataset — rows are observations, columns are variables. A list is a filing cabinet that can hold anything: text, numbers, plots, other lists, or even functions.
# VECTOR — R's most fundamental object marks <- c(88, 92, 76, 95, 81) # c() = combine/concatenate # Lab Q1 — sum, mean, product sum(marks) # 432 mean(marks) # 86.4 prod(marks) # product of all elements # Lab Q2 — airquality built-in dataset data(airquality) head(airquality) summary(airquality) colSums(is.na(airquality)) # count NAs per column # Lab Q3 — list of dataframes df_list <- list( batch1 = data.frame(name=c("Priya","Raj"), marks=c(88,74)), batch2 = data.frame(name=c("Zara","Arun"), marks=c(95,81)) ) # Lab Q4 — access each dataframe from list print(df_list[["batch1"]]) # by name print(df_list[[2]]) # by index # Lab Q5 — create, summarise, sort, add column, export students <- data.frame( name = c("Priya", "Raj", "Zara", "Arun"), marks = c(88, 74, 95, 81), passed = c(TRUE, TRUE, TRUE, TRUE) ) summary(students) # statistical summary str(students) # structure and data types students$grade <- ifelse(students$marks>=85, "A", "B") students_sorted <- students[order(-students$marks), ] # Export to Excel # install.packages("writexl") library(writexl) write_xlsx(students_sorted, "students_report.xlsx")
Session 17 · tidyverse & Data Manipulation
dplyr · ggplot2 · tidyr · rvest — the modern R ecosystem for data science
tidyverse is a LEGO system for data. Each package (dplyr, tidyr, ggplot2) is a compatible brick that snaps together. The pipe %>% connects them: "Take this data, THEN filter, THEN group, THEN summarise, THEN plot." You read it left to right like English — no nested parentheses, no intermediate variables to track.
library(tidyverse) # loads dplyr, ggplot2, tidyr, readr, etc. ═══ The 5 Core dplyr Verbs ══════════════════════════════════ diamonds %>% filter(cut == "Premium", price > 5000) %>% # 1. filter rows select(carat, cut, color, price) %>% # 2. pick columns mutate(ppc = price / carat) %>% # 3. add/transform col arrange(desc(ppc)) %>% # 4. sort head(10) # top 10 rows # group_by + summarise = SQL GROUP BY + aggregate diamonds %>% group_by(cut) %>% summarise( count = n(), avg_price = mean(price), max_p = max(price) ) %>% arrange(desc(avg_price)) ═══ Lab Q3: Pie Chart + Bar Chart ═══════════════════════════ # Bar chart ggplot(diamonds, aes(x=cut, fill=cut)) + geom_bar() + labs(title="Diamond Cut Distribution", x="Cut", y="Count") + theme_minimal() + theme(legend.position="none") # Pie chart (bar + coord_polar trick) cut_counts <- diamonds %>% count(cut) ggplot(cut_counts, aes(x="", y=n, fill=cut)) + geom_col(width=1, color="white") + coord_polar("y", start=0) + # this transforms bar → pie labs(title="Diamond Cut Proportions") + theme_void() ═══ Lab Q1: Load JSON and XML ════════════════════════════════ library(jsonlite); library(XML) json_df <- fromJSON("data.json", flatten=TRUE) xml_df <- xmlToDataFrame(xmlParse("data.xml")) summary(json_df); summary(xml_df) ═══ Lab Q2: Web Scraping with rvest ═════════════════════════ library(rvest) page <- read_html("https://books.toscrape.com") titles <- page %>% html_nodes("h3 a") %>% html_attr("title") prices <- page %>% html_nodes(".price_color") %>% html_text() books <- data.frame(title=titles, price=prices) head(books, 5)
Session 18 · R Functions & R Markdown
Writing reusable functions · ChickWeight case study · Reproducible reports
R Markdown is a lab notebook that runs its own experiments. You write prose, embed R code chunks, and when you click "Knit" in RStudio, every chunk executes and its output — tables, charts, statistics — appears right there in the final PDF or HTML. No copy-pasting from console. The report is the reproducible code.
# User-defined function calculate_gpa <- function(m1, m2, m3) { gpa <- (1/3)*m1 + (1/2)*m2 + (1/4)*m3 cat(sprintf("GPA = %.2f\n", gpa)) return(gpa) } calculate_gpa(80, 75, 90) # GPA = 79.17 ═══ ChickWeight Case Study ═══════════════════════════════════ data(ChickWeight) str(ChickWeight) # weight, Time, Chick, Diet # (a) Weight vs Time for Chick 34 chick34 <- ChickWeight[ChickWeight$Chick == 34, ] plot(chick34$Time, chick34$weight, type = "b", # b = both lines and points col = "#b34a22", pch = 16, # filled circles xlab = "Days after birth", ylab = "Weight (grams)", main = "Chick 34 — Growth Over Time") # (b) Boxplot for Diet group 4 diet4 <- subset(ChickWeight, Diet == 4) boxplot(weight ~ Time, data = diet4, col = "lightblue", main = "Diet 4: Weight by Time Point") # (c) Mean weight per time for Diet 4 mean_d4 <- tapply(diet4$weight, diet4$Time, mean) plot(names(mean_d4), mean_d4, type="b", col="#b34a22", pch=16, xlab="Time", ylab="Mean Weight (g)", main="Mean Weight — Diet 4 vs 2") # (d) Add Diet 2 line to same plot diet2 <- subset(ChickWeight, Diet == 2) mean_d2 <- tapply(diet2$weight, diet2$Time, mean) lines(names(mean_d2), mean_d2, col="#1a6e5c", type="b", pch=17) # (e) Legend and title legend("topleft", legend = c("Diet 4", "Diet 2"), col = c("#b34a22", "#1a6e5c"), lty = 1, pch = c(16,17), bty="n")
--- title: "PG-DBDA Data Analysis Report" author: "Your Name — C-DAC Kharghar" date: "`r Sys.Date()`" output: html_document: toc: true toc_float: true theme: flatly --- ## Introduction Analyse the **ChickWeight** dataset to compare dietary groups. ```{r setup, include=FALSE} knitr::opts_chunk$set(echo=TRUE, warning=FALSE, message=FALSE) library(tidyverse) ``` ```{r growth-chart, fig.width=9, fig.height=5} data(ChickWeight) ggplot(ChickWeight, aes(x=Time, y=weight, color=factor(Diet), group=Chick)) + geom_line(alpha=0.4) + stat_summary(aes(group=Diet), fun=mean, geom="line", linewidth=1.5) + labs(title="Chick Growth by Diet Group", color="Diet") + theme_minimal() ``` ## Conclusion Diet 4 shows the highest mean weight gain by day 21 (p < 0.05). # Inline R: "The dataset has `r nrow(ChickWeight)` observations."