• [SEACrowd] Public SEA Datasheets

    Available datasheet list can be accessed via https://tinyurl.com/SCDatasheets. Before filling this form, please kindly check whether the dataset is present in the list or not. Since this initiative is mainly to represent SEA, please make sure that the datasets are either collected: 1) from speakers in SEA or 2) in SEA regions. UPDATE: Submissions after 31 March 23:59 (UTC) will receive 0 points.
  • Public Datasheet

    All about the dataset hose datasets.
  • Dataset subset(s)
  • Dataset language(s)*
  • Dataset collection region — Where are the annotators from? Or is the data collected in/from specific SEA region(s)?*
  • Dataset task(s)*
  • Dataset modality*
  • Dataset domain(s)*
  • Dataset license*
  • Dataset annotation collection style*
  • Dataset annotation validation style*
  • Does the dataset have a pre-defined split?*
  • Please read this

    We understand that some datasets have a huge number of subsets and it will be very tedious to input the data sizes per subset one by one through the form. If the dataset has >10 subsets, please just input "Contact me for data size" for the data size questions, and we will help you inputting it.
  • Train split data size (per data subset)*
  • Validation split data size (per data subset)*
  • Test split data size (per data subset)*
  • No split data size (per data subset)*
  • Data size unit*
  • Does the data have PII (personally identifiable information)?*
  • Does the data have sensitive content/information?*
  • Dataset Credentials

  • Dataset provider/affiliation*
  • Publication venue*
  • Dataset access*
  • Contributor Information

    For clarity, we might want to do a follow-up about the details on your datasheet.
  • Are you...?*
    Rows

  • Should be Empty: