AI & Data Management
The Irving Institute for Cancer Dynamics (IICD) supports the responsible use of artificial intelligence (AI) to advance cancer research, strengthen administrative operations, and support academic activities. AI tools can increase efficiency and enable new analytical approaches, but they also introduce significant risks related to data privacy, confidentiality, intellectual property, and regulatory compliance.
All use of AI within IICD must comply with Columbia University policies governing AI, data protection, information security, and data retention. This page outlines expectations specific to IICD while pointing to authoritative University guidance.
Questions about appropriate AI use within IICD, data handling, or retention expectations should be raised before using AI tools. Requests for new AI tools or exceptions to policy must be reviewed through CUIT and the Office of the Provost. Noncompliant use of AI tools may result in corrective action.
The following Columbia University resources provide authoritative guidance governing AI use, data protection, privacy, and retention at IICD. These policies apply to all faculty, staff, researchers, trainees, and affiliates engaging in University-related work.
AI Policy
Data Protection, Privacy, and Retention
- https://www.cuit.columbia.edu/data-protection
- https://universitypolicies.columbia.edu/content/data-classification-policy
- https://research.columbia.edu/content/data-retention
- https://universitypolicies.columbia.edu/content/columbia-university-policy-records-retention
- https://universitypolicies.columbia.edu/content/email-usage-policy
- https://universitypolicies.columbia.edu/content/information-security-charter
Columbia University Information Technology (CUIT) has published University-wide best practices for responsible AI use. This guidance applies to faculty, staff, researchers, and students and reinforces a core principle: AI should assist human judgment, not replace it.
Key expectations emphasized by CUIT include:
- Protect University data and privacy at all times
- Use only CUIT-approved AI tools for sensitive or regulated information
- Verify the accuracy, relevance, and appropriateness of AI-generated output
- Maintain human oversight for decisions that affect people or institutional outcomes
- Monitor for bias and inequitable impacts
- Be transparent about when AI contributes to work products
The guidance is particularly relevant for administrative and research workflows, where improper AI use can introduce risks related to data exposure, compliance, and research integrity.
Best Practices for Responsible AI Use at Columbia University: https://etc.cuit.columbia.edu/news/best-practices-responsible-ai-use-columbia-university
AI tools may be used within IICD to support activities such as drafting and editing text, summarizing information, exploratory data analysis, coding assistance, and other administrative or research-support functions. AI may not replace required scientific judgment, human review, or formal decision-making processes.
All AI-assisted outputs must be reviewed for accuracy, bias, completeness, and appropriateness before being used in research, reporting, communications, or administrative actions. Responsibility for the final work product remains with the individual user.
AI tools must not be used to circumvent:
- Research integrity requirements
- Sponsored project or regulatory obligations
- Academic integrity standards
- Institutional approval or review processes
IICD handles highly sensitive research, personnel, and financial information. Unless explicitly approved through Columbia University Information Technology (CUIT) and central procurement, users must not input the following into generative AI tools:
- Personally identifiable information (PII)
- Protected health information (PHI)
- Confidential or restricted University information
- Unpublished research data, analyses, manuscripts, or grant proposals
- Proprietary, contractual, or third-party confidential materials
Entering sensitive or unpublished information into AI systems may result in loss of confidentiality, loss of intellectual property rights, violations of privacy law, or noncompliance with sponsor requirements.
AI use does not change existing expectations for responsible research conduct at IICD. Researchers remain fully responsible for:
- The accuracy and integrity of data and analyses
- Transparency regarding methods, including any use of AI tools
- Compliance with sponsor, journal, and regulatory requirements
AI tools must not be used to process unpublished research data or research subject information unless the tool has been explicitly approved for that purpose and appropriate safeguards are in place.
The use of AI tools does not alter Columbia University’s data retention requirements. Columbia distinguishes between research data and administrative records, each governed by different but complementary retention policies.
Research Data Retention
Research data generated or used within IICD must be retained for a minimum of three years after the end of a project, measured from the latest of final sponsor reporting, financial close-out, publication, or project cessation. Longer retention periods may be required for intellectual property protection, patenting activities, scientific misconduct reviews, or sponsor-specific requirements. Research data is owned by the University, with Principal Investigators serving as custodians responsible for its management and retention.
Administrative Records Retention
Administrative records, including finance, human resources, operational, and student-related records, are governed by the Columbia University Policy on Records Retention and associated schedules published in the Policy Library. Retention periods vary by record type and may range from several years to permanent retention, particularly for certain student education records or materials subject to archival requirements.
AI Outputs and Records Retention
AI tools must not be treated as systems of record. When AI-generated outputs are relied upon for official purposes—such as research conclusions, sponsor reporting, financial decisions, personnel actions, or formal communications—they may constitute University records and must be retained in accordance with the applicable research or administrative retention schedule. Exploratory or draft AI outputs that are not relied upon for official decision-making do not require retention; however, all required source data, final analyses, and official records must be preserved in approved University systems.
The use of AI tools does not alter Columbia University’s data retention requirements.
Research data generated or used within IICD must be retained in accordance with sponsor requirements and University policy, generally for a minimum of three years after the completion of a project, and longer where required. Administrative and financial records must be retained according to applicable University records retention schedules.
AI platforms must not be treated as systems of record. Users are responsible for ensuring that:
- Original research data and required documentation are stored in approved University systems
- AI-generated outputs that constitute official records are retained appropriately
- Data is not lost, altered, or improperly disposed of through reliance on AI tools
Columbia University Information Technology provides University-approved AI tools that operate within managed environments designed to protect institutional data. IICD staff and researchers are expected to use these tools, rather than consumer or personal AI accounts, when working with University-related materials.
Columbia AI
University-Supported AI Tools
- CUIT AI Services – https://www.cuit.columbia.edu/content/ai-services
Central listing of University-supported AI platforms, services, and consultation-based offerings reviewed for security, privacy, and compliance. - ChatGPT Education – https://www.cuit.columbia.edu/content/chatgpt-education University-licensed access to OpenAI’s ChatGPT with enhanced security and full functionality, including advanced data analysis and image generation
- CHAT – https://www.cuit.columbia.edu/www.cuit.columbia.edu/content/chat Columbia’s customizable AI chat platform, powered by LibreChat, supporting multiple models in a University-managed environment
- Google Gemini – https://www.cuit.columbia.edu/content/google-gemini Multimodal AI available through LionMail for text, image, and code tasks
- NotebookLM – https://www.cuit.columbia.edu/content/notebooklm Google’s AI research assistant integrated with Google Workspace for document synthesis and knowledge management
Consultation-Based AI Services
CUIT also offers AI-enabled services for approved research or administrative use, including:
- Audio transcription
- Text anonymization to remove personally identifiable information
- Automated text mining of large document collections
These services are available by consultation to ensure appropriate data handling and compliance. Please email [email protected] for more information. Email [email protected] for support.
