Section outline

    • DIMER: Explore, Experiment, Evaluate

      The activity is modular and self-paced. The four levels suggest a progression; individual notebooks can be taken independently. Learners should have basic Python and Colab familiarity, but need no prior ML experience.

      Explore how AI models classify data, predict numerical values, forecast future observations, retrieve information, and interpret text, audio, images and documents. Each guided notebook helps you understand a task, run a baseline and model comparison, conduct a controlled experiment, and explain the results using evidence.

      You do not need to understand every line of code or the underlying model architecture. Begin with the relationship:

      Input → Model/System → Output

      Some notebooks compare individual models; others combine several stages into a system. Your task is to understand what goes in, what happens at each stage, what comes out, and how to judge the result.

      Questions or support needs? Post in the course’s Office Hours Session — Live Q&A and Course Consultation. Selected questions will be addressed during the session. Share error messages only after removing personal information, credentials and sensitive data.

      Learning objectives

      By the end of a selected learning unit, you should be able to:

      • Identify the AI task, expected inputs and resulting outputs.
      • Run a guided notebook and explain the purpose of its baseline.
      • Interpret evaluation metrics alongside example successes and failures.
      • Predict the effect of changing one approved input or setting, then compare outcomes.
      • Draw a conclusion supported by observations, including limitations.
      • Recognise data-permission, privacy, bias and human-review considerations.
      • Describe a realistic application and the evidence needed before practical use.

      Before you begin

      You will need:

      • Access to Google Colab.
      • A DIMER training account for exploring the model repository.
      • Basic familiarity with Python and running notebook cells.
      • A notebook selected from the catalogue below.
      • A Hugging Face account or access token only where required by your selected notebook.

      Open the notebook using its Open in Colab badge. Read its runtime requirements and select the recommended hardware before execution. Start with the supplied dataset and default settings.

      Leave installation, model-loading, dataset-split and evaluation setup cells unchanged unless the notebook explicitly directs you to modify them.

      Change runtime type T4 GPU → Connect → Run all

      Access and participation: If account access, GPU availability, bandwidth, cost or accessibility needs prevent execution, contact the facilitator for a supported alternative, such as a demonstration or review of supplied outputs. Ask for guidance before purchasing compute or a subscription. Where a notebook includes audio, use its transcripts, captions and visualisations where available; request support if these do not provide an accessible equivalent.

      Hugging Face access — only when required

      Follow the selected notebook’s authentication instructions. If it supports anonymous access, you do not need to create or grant access to a token.

      Where authentication is required:

      1. Create or sign in to your Hugging Face account.
      2. Visit the required model or dataset page. Complete any access request or licence acceptance using that account.
      3. Open Settings → Access Tokens and create a token for this activity.
      4. Prefer a fine-grained token with read access limited to the required resources. Write access is not needed for this activity.
      5. In Colab, open Secrets using the key icon. Add the secret name specified by the notebook, commonly HF_TOKEN.
      6. Enable Notebook access only for notebooks that require and use that secret.

      Creating a token does not itself grant access to gated resources.

      Keep your token private. Do not paste it into notebook code, outputs, screenshots, forum posts or learning records. If exposed, revoke it immediately and create a replacement. Remove notebook access when it is no longer needed.

      See the Hugging Face access-token guide for permissions and token management.

      Responsible exploration

      These notebooks support learning and experimentation. Successful execution does not establish reliability, fairness or readiness for deployment. Some notebooks remain Candidate pending further validation and evidence review.

      • Use permitted data. Follow the privacy and BYOD guidance below. Public availability does not automatically establish permission to upload or reuse data.
      • Check important outputs. Compare predictions with appropriate references. Fluent text, plausible images, high scores and convincing visualisations can still be wrong.
      • Consider representation and uneven performance. Ask which languages, accents, locations, categories or input conditions are missing from the sample, and who could be affected by errors.
      • Keep people responsible for decisions. Do not use activity outputs directly to determine consequential decisions about people or safety. Practical use requires appropriate validation, accountable human review and applicable institutional approval.
      • Respect licences and attribution. Follow the terms for code, models, datasets and media. Download access does not imply unrestricted reuse, redistribution or commercial use.
      • Inspect before sharing. Notebook outputs, screenshots, logs and exported files may contain sensitive information even when the code does not.
      • Report concerns. Stop the affected activity and contact the facilitator if you encounter exposed data, harmful outputs or a suspected security issue. Do not reproduce sensitive material in a public forum.

      Choose your learning unit

      The catalogue contains 32 learning entries, including an optional continuation on image representations.

      • Foundations: Learn predictions, metrics, baselines, held-out evaluation, representations and retrieval.
      • Applied tasks: Apply the same evaluation habits to different inputs and outputs.
      • Composed systems: Trace how several stages work together and identify where errors enter the system.
      • Scientific applications: Apply earlier lessons to spatial, temporal and domain-shift problems.

      Level Notebook/Learning Unit Model/s Link to Colab Use Case / Description
      1. Foundations FreshRetailNet Multi-Model Classification v2 Mitra, TabDPT v1.2, TabPFN-3, TabICLv2 Open In Colab Predict retail sales categories and compare models against baselines using held-out classification metrics.
        FreshRetailNet Multi-Model Regression v2 Mitra, TabDPT v1.2, TabPFN-3, TabICLv2 Open In Colab Predict numerical retail outcomes and compare prediction errors across models and baselines.
        Multi-Model Time-Series Forecasting TiRex-2, Chronos-2; Toto 2.0 in FULL mode Open In Colab Forecast future observations while respecting time order, forecast horizons and independent evaluation periods.
        Multi-Model Image Classification MobileNetV4, ResNet-50, ConvNeXt-Tiny, ViT-B/16; SwinV2-Tiny and EVA-02 in FULL mode; optional DINOv2 Open In Colab Compare pretrained image representations through a shared classification task and controlled evaluation.
        Modern Image Classification — Representations and Transfer Learning ResNet-50, MobileNetV4, ConvNeXt-Tiny, ViT-B/16, SwinV2-Tiny, EVA-02 Open In Colab Study representations, 5-NN, PCA, data efficiency and transfer learning after the introductory vision unit.
        Qwen3 Semantic Search and Reranking Qwen3-Embedding-0.6B, Qwen3-Reranker-0.6B Open In Colab Retrieve relevant text using embeddings and examine how reranking changes the shortlist.
      2. Applied tasks Document-Type Classification DiT Base fine-tuned on RVL-CDIP Open In Colab Classify document images and investigate category confusion and sensitivity to image transformations.
        OCR and Structured Document Extraction GOT-OCR 2.0, SmolDocling Open In Colab Extract text and document structure; inspect transcription errors, reading order and structured outputs.
        Multi-Model Closed-Set Object Detection RT-DETR R50-VD, YOLOX-S; YOLOX-X in FULL mode Open In Colab Locate objects from a fixed category vocabulary and compare detection quality and compute tradeoffs.
        Open-Vocabulary Object Detection Grounding DINO Tiny, OWLv2 Base/16 Ensemble Open In Colab Locate objects using text descriptions and explore sensitivity to prompts and thresholds.
        Multi-Model Image Matching LightGlue + ALIKED, XoFTR Open In Colab Find correspondences between images and assess whether matches support reliable geometric alignment.
        Vision-Language Retrieval SigLIP 2, SigLIP v1, BLIP ITC and ITM Open In Colab Retrieve images and captions, compare embedding-based retrieval and inspect the effect of reranking.
        Whisper Speech Recognition Whisper large-v3-turbo Open In Colab Transcribe speech, measure word and character errors, and examine the effect of background noise.
        Multi-Model Image Segmentation CLIPSeg, SAM, SAM 2 Open In Colab Produce pixel masks from prompts, adapt CLIPSeg and compare SAM models under matched geometric prompts.
        Multi-Model Depth Estimation Depth Anything V2 Small, ZoeDepth Open In Colab Distinguish relative depth from metric distance and evaluate predictions against sensor references.
        Sound Event Classification Audio Spectrogram Transformer (AST), with an adapted ESC-10 classification head Open In Colab Recognise environmental sounds, compare initial and trained heads, and investigate sensitivity to added noise.
        Zero-Shot Text Classification — Do Our Labels Change the Answer? BART-large-MNLI Open In Colab Classify text using category descriptions, compare pretrained and adapted models, and test sensitivity to label wording.
        English–Tagalog Translation MarianMT / OPUS-MT-en-tl Open In Colab Translate English into Tagalog and assess meaning preservation alongside automatic translation metrics.
        Text Summarization BART-large-CNN Open In Colab Produce shorter versions of source texts and examine information retention, omissions and factual fidelity.
      3. Composed systems Document QA LayoutLM Document QA, Pix2Struct DocVQA Open In Colab Compare answers extracted from OCR and layout with answers generated directly from document images.
        Open-Vocabulary Detection and Counting Grounding DINO Tiny, OWLv2, CountGD Open In Colab Compare detection-based counting with text, exemplar and combined counting prompts.
        Table Intelligence Table Transformer Detection, Table Transformer Structure Recognition, TAPAS Large WTQ Open In Colab Locate tables, reconstruct their structure and answer questions while tracing errors between stages.
      4. Scientific applications AI for Earth Observation and Climate Applications Prithvi-EO-2.0-300M; Prithvi flood, burn-scar and crop-classification models Open In Colab Explore multispectral satellite representations and applications in flood mapping, burn-scar detection and crop classification.
        Weather and Earth-System Forecasting Aurora Small, with bounded LoRA adaptation Open In Colab Forecast atmospheric and surface variables; compare persistence, frozen-model and adapted forecasts under resolution shift.
        Philippines Flood Mapping — Mapping Water After Typhoon Ompong Prithvi-EO-2.0-300M-TL-Sen1Floods11; MNDWI and no-water baselines Open In Colab Map clear-sky surface water in Candon and Vigan, Ilocos Sur, after Typhoon Ompong. Compare frozen Prithvi with spectral and simple baselines, evaluate held-out geography, and distinguish surface water from confirmed flood inundation.
        Philippine Biodiversity Field-Survey Triage — Mapping Similarity, Not Certainty BioCLIP 2, SigLIP 2; nearest-neighbour and fitted linear classifiers Open In Colab Compare general and biological image representations for four Philippine endemic bird species. Evaluate photographs from held-out observers, explore common versus scientific names, and refer uncertain predictions for expert review without treating model scores as confirmed species identification.
        Philippine Dengue Forecasting — Do Weather Signals Improve Predictions? Chronos-2, Mitra Regressor with and without weather; persistence, seasonal naïve and Ridge baselines Open In Colab Forecast Quezon City’s reported dengue counts one to four reporting blocks ahead. Test whether historical weather improves predictions, examine reporting-delay assumptions and interval coverage, and distinguish exploratory results from operational health warnings. 
        Philippine Rice Pest Surveillance — From Insect Annotations to Counting DETR, BioCLIP 2 with an adapted species head, SigLIP 2; classification and counting baselines Open In Colab Detect, classify and count rice pests using published image annotations. Trace missed detections and species errors through the pipeline, flag uncertain predictions for review, and distinguish exploratory benchmark results from validated field surveillance.
        Filipino Audio Archive Search — How Do Transcription Errors Affect What We Find? Whisper large-v3-turbo, Qwen3-Embedding-0.6B, Qwen3-Reranker-0.6B; BM25 baseline Open In Colab  Search Filipino speech recordings through automatic transcripts. Compare keyword, semantic and reranked search; trace transcription errors into retrieval failures; and explore candidate depth. Current retrieval results are engineering diagnostics pending human relevance review.
        Synthetic Defect Augmentation — Can Generated Images Improve Visual Inspection? Phi-4-multimodal-instruct, FLUX.1 [schnell], frozen ResNet-18 with fitted linear classifiers Open In Colab Generate candidate scratch and spot images from descriptions of Bosch surface-defect examples. Compare real-only, conventional and synthetic augmentation under matched training budgets, evaluate on real held-out photographs, and distinguish visual plausibility from measured usefulness.
        Philippine Reef Heat-Stress Outlook Chronos-2; persistence and seasonal-reference baselines Open In Colab Forecast regional marine HotSpot indicators at 7, 14 and 28 days and calculate accumulated Degree Heating Weeks. Compare chronological forecast errors, explore history length, and distinguish thermal-stress indicators from observed coral bleaching or operational warnings.
        Small-Business Receipt Intelligence — From Receipts to Records Tesseract OCR; frozen and receipt-adapted LayoutLM; keyword-rule baseline Open In Colab Extract totals, subtotals, tax and service charges from Indonesian receipt photographs. Trace OCR and extraction errors, compare rules with pretrained and adapted models, and evaluate a human-review referral policy without treating unflagged amounts as verified accounting records.

      Catalogue notes:

      • Modern Image Classification is an optional continuation after the introductory image-classification unit.
      • Closed-set detection provides a useful reference before open-vocabulary detection.
      • The model column lists principal models; notebooks may also include simple or classical baselines.
      • Optional and FULL-mode models do not necessarily run under the default settings. Follow the selected notebook’s resource guidance.

      Activity procedure

      A. Understand the task and make a prediction

      Read the notebook’s introduction and identify:

      • The task and model or system being explored.
      • The input format and expected output.
      • The baseline used for comparison.
      • The metric or observation used to judge performance.

      Before running the example, write a short prediction:

      What do you expect to happen, and why?

      B. Run and interpret the default example

      Run the notebook as instructed, using its default configuration.

      Examine the baseline, model results, visual examples and evaluation summaries. Use the notebook’s “what to notice” guidance to interpret them.

      Consider:

      • How does the model compare with the baseline?
      • Which examples succeed or fail?
      • What does the metric measure—and what does it leave out?
      • Do the displayed examples support the overall results?
      • Could an apparently confident or plausible output still be incorrect?

      A notebook completing successfully confirms execution. It does not, by itself, establish that the model performs well.

      C. Change one thing and compare

      Use the notebook’s designated learner activity. Predict the effect of the change before running it.

      Depending on the learning unit, the activity may change an input example, prompt, retrieval setting or another explicitly identified parameter. Keep the other settings fixed so the comparison remains interpretable.

      Record the original setting, changed setting and observed result. Explain whether the outcome supported your prediction.

      Preserve the notebook’s training, validation and test roles. Use validation data for tuning where instructed. If you explore changes using test results, label the exercise as exploratory; those same results are no longer independent evidence for the selected setting.

      D. Bring your own data — optional

      Attempt BYOD only after completing the supplied example and reading the notebook’s BYOD instructions.

      Prepare a small, representative dataset that meets the stated requirements. These may include supported file types, required columns, labels, image dimensions, paired inputs, minimum sample counts or separate training, validation and test sets.

      Run the provided validation checks before model execution. If validation fails, correct the data or return to the supplied example; do not bypass the checks.

      Data safety and privacy:

      • Use only data permitted for upload to the selected platform.
      • Do not upload confidential, classified, proprietary, personally identifiable or otherwise restricted information.
      • Apply these checks to text, documents, images, voices, location information and file metadata.
      • Public availability or de-identification alone does not establish permission.
      • Do not collect sensitive personal data simply to complete a fairness or bias exercise.
      • If permission is uncertain, use the supplied sample data.
      • Before sharing a notebook or export, inspect its saved outputs, logs, filenames and metadata.
      • Remove activity uploads and saved copies you no longer need, following applicable retention requirements. Ending a runtime does not remove copies saved elsewhere.

      Record what you supplied, the output produced and how you assessed it. Without suitable reference labels or measurements, describe the result as a qualitative observation rather than an accuracy claim.

      E. Conclude with evidence and consider an application

      Summarise what the experiment showed, what it did not establish and one limitation that matters.

      Then connect the capability to a realistic problem:

      • What problem could it help address?
      • What data and reference answers would be needed?
      • Which languages, locations, categories or input conditions were missing or poorly represented?
      • Could errors affect some people or situations more than others?
      • What harm could an incorrect result cause?
      • Who would review the output, and when should the system’s recommendation be rejected?
      • What further testing and approval would be needed before practical use?

      You may use this conclusion scaffold:

      On [sample and task], the model produced [observed result] compared with [baseline]. Changing [one factor] led to [observed difference]. This supports [limited conclusion], but does not establish [broader claim]. Before practical use, we would need [additional evidence and human review].

      Completion records are optional personal learning notes, not required submissions.

      F. Continue exploring and apply your learning

      Connect what you learned to the model pipelines available on DIMER. Browse the DIMER model repository, choose a pipeline relevant to a realistic problem, and explore its end-to-end notebook.

      These pipeline notebooks provide a deeper look at data preparation, model adaptation where supported, evaluation and inference.

      Define the expected inputs, useful outputs and what would count as success. Using suitable data that you have permission to upload, compare a baseline with model results, inspect failures and identify where human review would be needed.

      Conclude with an evidence-based recommendation: proceed with further testing, revise the approach, or stop. Explain what your results support, their limitations and what additional evidence would be needed before practical use.

      Help inform the future of DIMER

      Separately from the learning activity and Exit Survey, you are invited to consider participating in the AIaaS Market Study, which explores organisations’ AI needs, readiness and potential applications.

      Participation is voluntary and is not a requirement for this learning activity. Review the study’s information, eligibility requirements, privacy notice and consent process before deciding. Respond in a personal or organisational capacity only as permitted by the study and your organisation; do not disclose confidential organisational information.

      → Participate in the DIMER Market Study survey