Handmade Datasets: Strategies for working critically with small data and Artificial Intelligence
.・゜゜・*:・゚✧ ₊ ⊹ . ˖ . .・゜゜・..⋆。⋆˚⋆.・.・☾.・゜゜
Handmade Datasets is currently running as a course taught by Aarati Akkapeddi at IDM, NYU.
Class Hours: Mondays 5:00 – 7:50pm | Room 310
Instructor email: aa9570[at]nyu.edu
₊˚ ✧---. ˖--・--------------⊱⋆⊰--------------・--˖ .--- ✧ ˚₊
This course unpacks the data pipeline behind large AI systems. Students will examine how websites get crawled, how images and their alt-text descriptions become training material, how click-workers and automated systems process data, and how all of this shapes what a model generates. The course will propose an alternative to these systems: "handmade datasets," or human-scale, personally assembled datasets. Students will learn practical techniques for training smaller kinds of AI models with limited data. Through hands-on projects and critical discussion, students will develop both technical capability and a more nuanced understanding of data collection ethics and the labor behind AI systems. By the end of the course, participants will have created their own handmade dataset and trained a custom model.
No prior machine learning or advanced coding experience required. Students should be comfortable with basic computer literacy (file management, running applications) and have a willingness to experiment with new tools. We will focus on understanding concepts, collecting meaningful data, and remixing existing code rather than writing algorithms from scratch. An interest in questions around data ethics, labor, representation, personal archives, or critical approaches to technology is encouraged.
Students will complete weekly assignments to experiment with generative text, image, and audio models, expanding on one to work with as a final project. Participants can expect to spend ~5 hours weekly on work outside of class.
By the end of this course, students will:
- Create a handmade dataset and train a custom model
- Understand the data pipeline of text-to-image AI systems, from web scraping through training and recognize how dataset composition shapes model outputs and reproduces biases
- Gain hands-on experience with multiple approaches to models with small datasets.
- Learn methods for collecting, organizing, and augmenting personal datasets.
- Understand the difference between fine-tuning or transfer learning vs. training from scratch
- Develop insight into the hidden labor behind large-scale AI and how to center consent, ownership, and stewardship when working with data.
- Learn to troubleshoot common technical challenges (mode collapse, training instability, computational limits, dataset balance) and make informed decisions about model selection based on available data training resources, and ethical & creative considerations
₊˚ ✧-. ˖—----------------------・-⊱⋆⊰-・----------------------˖ .- ✧ ˚₊
This course is taught on Lenape Land: The Lenni Lenape peoples are the original and continued rightful stewards of the land that I love and on which I work and reside. I encourage everyone to learn about the history of the land they live and work on.
₊˚ ✧-. ˖—----------------------・-⊱⋆⊰-・----------------------˖ .- ✧ ˚₊
Course Requirements
Attendance
Students are expected to attend classes in-person regularly. This class covers a wide scope and each week is important. Two absences are allowed; if you need to miss any other class you may seek the appropriate documentation through the Student Affairs office. Contact Deanna Rayment deanna.rayment@nyu.edu, the Coordinator for Student Advocacy and Compliance. Any additional absence will cause your grade to be reduced by 1/2 a grade point (e.g. A to an A-).
Be on Time. If you are late two times, I will mark you as absent.
Late Assignment Policy
I will not accept late assignments. Many of our assignments build off of each other and you will quickly fall behind if you miss deadlines. If you anticipate missing an assignment deadline due to an emergency or illness please contact me as soon as you can.
Class Participation
Attendance alone does not guarantee an A in class participation. This course includes discussion and group critique. Engaged and motivated participation from each student is necessary in creating an environment where everyone can get the most out of this course.
Handmade Dataset Experiments
While these projects could influence your final project idea, the first three model experiments are intentionally constrained. You are not expected to develop a fully fleshed out creative concept for each one. The goal is to rapidly encounter several different approaches to machine learning and develop an intuition for how they behave with small datasets. And then your final project is where you will have the time to fully develop something you care about.
- Image
- Text
- Audio
Handmade Dataset Final Project
You will create a project training a model on your own handmade dataset. In class we will cover models that generate images (GAN, LoRA), text (Ollama + RAG, Microgpt), and sound (RAVE). You can use any of the methods or a combination of methods we covered thus far in class for your final project. You are welcome to create work in the realm of art, design, or tool. Because building handmade datasets and training takes time, labor and a lot of trial and error. Your first dataset may not work the way you expect. Your model may produce unexpected results. You may discover that your dataset needs to change, that you need more or different data, or that another training approach makes more sense. Therefore, we will spend the majority of the semester on this one final handmade dataset project. We will guide you through this project in four milestones:
- Milestone 1 [Week 9]: Present Final Concept, datasheet and data samples
- Milestone 2 [Weeks 11]: Present dataset (should be at least 70% complete)
- Milestone 3 [Week 13]: Training Progress Check-in
- Milestone 4 [Week 14-15]: Explain the project, dataset, model, results, iterations, and discoveries; present final work at showcase
Reading Recipes
These assignments are meant to inspire and strengthen your understanding of data provenance, data labor and other ethical and creative issues related to generative AI and strategies for working with small data.
- Karen Hao - Research a Handmade Dataset project
- Allison Parrish - Markov Chain Activity
- Moisés Horta Valenzuela - Listening Activity
- Suzanne Kite - Reflection Diagram
Grading of Assignments
The grade for this course will be determined according to the following:
| Assignments/Activities | % of Final Grade |
|---|---|
| Class Participation | 16% |
Handmade Datasets Experiments
|
24% |
Reading Recipes
|
20% |
Handmade Dataset Final Project
|
40% |
Letter Grades
Letter grades for the entire course will be assigned as follows:
| Letter Grade | Points | Percent |
|---|---|---|
| A | 4.00 | 92.5% and higher |
| A- | 3.67 | 90.0% – 92.49% |
| B+ | 3.33 | 87.5% – 89.99% |
| B | 3.00 | 82.5% – 87.49% |
| B- | 2.67 | 80.0% – 82.49% |
| C+ | 2.33 | 77.5% – 79.99% |
| C | 2.00 | 72.5% – 77.49% |
| C- | 1.67 | 70.0% – 72.49% |
| D+ | 1.33 | 67.5% – 69.99% |
| D | 1.00 | 62.5% – 67.49% |
| D- | 0.67 | 60.0% – 62.49% |
| F | 0.00 | 59.99% and lower |
Course Schedule [subject to change]
| Week/Date | Topic | Homework |
|---|---|---|
| Week 1 9/14/26 |
Course logistics and introductions • The Data Pipeline Behind Large AI Systems: Web scraping, labeling, click work, automated data collection • What is a “handmade dataset”? What is “small” in the context of AI? |
Recipe #1 Complete Intake form InstallGit and VSCode if you don’t already have them installed on your comptuer |
| Week 2 9/21/26 |
Introduction to GANs • What can a GAN learn from a small dataset? |
Handmade Dataset Experiment #1 Part I |
| Week 3 9/28/26 |
Data augmentation with images • GAN training |
Handmade Dataset Experiment #1 Part II |
| Week 4 10/5/26 |
Present Handmade Dataset Experiment #1 • Intro to Generative Text -> Markov Chains & LSTM |
Reading Recipe #2 Handmade Dataset Experiment #2 Part I |
| Week 5 10/12/26 |
Intro to Small/local Language Models RAG | Handmade Dataset Experiment #2 Part II |
| Week 6 10/14/26* |
Guest Speaker TBD • Present Generative Text |
Handmade Dataset Experiment #3 Part I Reading Recipe #3 |
| Week 7 10/19/26 |
Introduction to Generative Audio models • Training Audio Models |
Handmade Dataset Experiment #3 Part II |
| Week 8 10/26/26 |
Present Audio models • Official Introduction to Final Project |
Work on Final Concept Presentation (~15 min to present, 5 min for feedback) |
| Week 9 11/2/26 |
Milestone 1: Present Final Concept, datasheet and data samples | Work on Datasheet |
| Week 10 11/9/26 |
Augmentation clinic | Work on dataset |
| Week 11 11/16/26 |
Milestone 2: Present Dataset | Work on dataset and training |
| Week 12 11/23/26 |
Training clinic Integrating models on the web dem |
Work on training/retraining if necessary |
| Week 13 11/30/26 |
Milestone 3: Training Progress Check-in • Integrating models on the web demo continued |
Work on Final Presentation |
| Week 14 12/7/26 |
Final Presentations | Work on Final Presentation Reading Recipe #4 |
| Week 15 12/14/26 |
Milestone 4: Final Showcase | :-) |
Course Materials
All readings will be provided. For certain techniques you will need access to GPU that may exceed the capabilities of your laptop. We will go over what NYU resources there are for this. Alternatively, or perhaps in the future beyond your time at NYU, you may opt for cloud services like Google Colab. We’ll discuss the costs associated with this.
Resources
- Access your course materials: https://handmadedatasets.com
- Databases, journal articles, and more: Dibner Library (library.nyu.edu)
- Assistance with strengthening your writing: NYU Writing Center (nyu.mywconline.com)
- Obtain 24/7 technology assistance: IT Service Desk (NYU IT) (nyu.edu/it/servicedesk)
Academic Honesty/Plagiarism
NYU School of Engineering Policies and Procedures on Academic Misconduct (from the School of Engineering Student Code of Conduct)
Introduction: The School of Engineering encourages academic excellence in an environment that promotes honesty, integrity, and fairness, and students at the School of Engineering are expected to exhibit those qualities in their academic work. It is through the process of submitting their own work and receiving honest feedback on that work that students may progress academically. Any act of academic dishonesty is seen as an attack upon the School and will not be tolerated. Furthermore, those who breach the School’s rules on academic integrity will be sanctioned under this Policy. Students are responsible for familiarizing themselves with the School’s Policy on Academic Misconduct.
Definition: Academic dishonesty may include misrepresentation, deception, dishonesty, or any act of falsification committed by a student to influence a grade or other academic evaluation. Academic dishonesty also includes intentionally damaging the academic work of others or assisting other students in acts of dishonesty. Common examples of academically dishonest behavior include, but are not limited to, the following:
- Cheating: intentionally using or attempting to use unauthorized notes, books, electronic media, or electronic communications in an exam; talking with fellow students or looking at another person’s work during an exam; submitting work prepared in advance for an in-class examination; having someone take an exam for you or taking an exam for someone else; violating other rules governing the administration of examinations.
- Fabrication:including but not limited to, falsifying experimental data and/or citations.
- Plagiarism: intentionally or knowingly representing the words or ideas of another as one’s own in any academic exercise; failure to attribute direct quotations, paraphrases, or borrowed facts or information.
- Unauthorized collaboration:working together on work that was meant to be done individually.
- Duplicating work:presenting for grading the same work for more than one project or in more than one class, unless express and prior permission has been received from the course instructor(s) or research adviser involved.
- Forgery:altering any academic document, including, but not limited to, academic records, admissions materials, or medical excuses.
Access the entire School of Engineering Student Code of Conduct here: engineering.nyu.edu/academics/code-of-conduct
Generative Tool Use In This Class
In this particular course there are a few different kinds of work you’ll be doing and I have slightly different policies for each:
- Ideating
Please don’t use LLMs (Large Language Models like chatGPT, Gemini, Claude, CoPilot etc) and/or text-to-image generation models (MidJourney, Stable Diffusion, Nano Banana, etc) to come up with ideas for the work you do in this course. Believe in yourself. If you are stuck, take a walk. Inspiration is all around you. Plan to gift yourself as much time and space as you can so that you don’t have to rush your assignments. - Assembling your own dataset to train an ML model
Unless explicitly mentioned in assignment directions, LLMs (chatGPT, Gemini, Claude, CoPilot etc) and text-to-image generation models (MidJourney, Stable Diffusion, Nano Banana, etc) are not permitted in the creation of your dataset. The idea of a handmade dataset is to work slowly and intentionally collecting/creating your own data, considering ownership, privacy, and labor along the way. - This course will also include readings/listenings.
None of the (four) readings I assign will be very long or jargony. You should not use LLMs to summarize readings. You are a unique person and you may resonate with something in a reading that the LLM glosses over in its summary. I don’t want you to lose out on that chance to connect with the texts/listenings. - Interacting with python code to Process/Augment Datasets and train models
I will be providing you with python scripts that you can remix. Occasionally for certain utilities or customizations that are unique to your specific project it is permissible to use an LLM to assist you on one condition: please document the changes that the LLM made. I ask for this because otherwise it can be really difficult for me to assist you in troubleshooting your code if we are both in the dark about what changes were made to the original file. I recommend double and even triple checking your file and folder paths before turning to an LLM to help with errors. - Integrating custom trained models into web applications/other coded interfaces
You may want to integrate the custom models you build into other things. For example, maybe you build a web application with your image-generating model so others can create images with it. In this case, it is permissible in this class to use LLMs to aid in coding/prototyping. I would still encourage you to lean on your own design skills and make anything you build with an LLM your own (for example, be intentional about aesthetic, fonts, instructional text, user flow, etc as opposed to just accepting what’s generated for you). And similarly to the above, it always helps me to help you if you keep track of what changes the LLM is making.
* Other permissible uses for LLMs would include assistance due to a disability (for example, using an LLM to describe visuals in place of a traditional screen reader) or to help with translation (for example, you may be most comfortable writing in your native language first and then translating to English).
Academic Accommodations
If you are a student with a disability who is requesting accommodations, please contact New York University’s Moses Center for Student Accessibility (CSA) at 212-998-4980 or mosescsa@nyu.edu. You must be registered with CSA to receive accommodations. Information about the Moses Center can be found at https://www.nyu.edu/csa. The Moses Center is located at 726 Broadway on the 2nd floor.
If you are experiencing an illness or any other situation that might affect your academic performance in a class, please email the Office of Advocacy, Compliance and Student Affairs: eng.studentadvocate@nyu.edu.
Statement On Inclusion
The NYU Tandon School values an inclusive and equitable environment for all our students. I hope to foster a sense of community in this class and consider it a place where individuals of all backgrounds, beliefs, ethnicities, national origins, gender identities, sexual orientations, religious and political affiliations, and abilities will be treated with respect. It is my intent that all students’ learning needs be addressed, and that the diversity that students bring to this class be viewed as a resource, strength and benefit. If this standard is not being upheld, please feel free to speak with me.
Resources For Non-Citizen Students
More than 40 percent of NYU students are international students. A smaller number are undocumented students, but many more come from mixed status families and communities. As a professor, I am committed to doing everything I can to ensure that every student, regardless of immigration status, is safe in this classroom. Following the recommendation of the NYU chapter of the AAUP, I encourage students to seek free legal support and other resources through NYU’s Immigrant Defense Initiative. NYU IDI provides an extensive list of updates and resources. Students may also consult the “Know Your Rights” information provided by the New York Immigration Coalition.