01 ProjectsHAI ProjectOverview Draft

Prompt to claude `Ok so right, the as far as I understand, the core idea is to help users gain trust in the AI agents, llms, or systems they use in their workflow so that better productivity coupled with reliability and reproducibility is achieved. The problems that corporations or research labs face is the stakes that come with the ceos or employees or managers using those AI helpers and blindly trusting or not trusting them come to. I remember the CEO of a bi gaming company Tencent ig, when he didnt want to pay the developers of Subnautica 2 for their success and turnover on their approved deal and the metric they had mutually decided, went to ask AI about how to solve this case, and AI just straight up suggested him to fire them, leading to many problems later on with the company's image suffering too. So our idea is to collaborate with corporations and when they use ai to generate any business solution or in case of research labs to generate any results or researchs, our confidence checker will analyse the output generate on various trust metrics and different perspectives which it may have missed as well, on benchmarks of real world previous cases, or actual physical data, and then generate such outputs be it visuals or text or something else to enable the user to trust the output of the ai agent or need to revise it. This will enable us to flag all the stakes to a company and not let them take business decisions which will adversely affect them like missing out on some important documents or for example a case when an auditing company arthur anderson audited another company's ledgers and missed out on many shell comapnies cashflow leading to disastrous losses. We will not be stopping on just confidence scores only, we also need to give actionable advice on how or what steps the consumer can take to increase the confidence score, and consquently gnerate reliable and trustworthy insights, or for example what decision he can take to solve his problem. We have a vision to first make this project research paper worthy that can be presented at ACM CHI Student conference, as well as other conferences, and then flesh it out and evolve it into a full fledged startup that caters to businesses and research institutions, since they have an incentive to produce reliable final outputs, as compared to daily users who if on using our product start trusting ai outputs less, will hamper the usage of it and make us an enemy in the eyes of these ai big tech corporations. Now that I have won the JAW Appcair Hackathon organised by Gikagraph.ai, we are well accquainted with the founder of the company and looking at it properly we can see that the work they are doing is very closely related to our project. Its like they are enabling users to trust the financial audtiting documents usign ai insights and grpahs, etc. So i want you to analyse them as a very good potential stakeholder for this project, where we can contact them to provide us funding or real world data which we can work upon, or soemthing else by engaging in a partnership with them. Below attached are the extra citations, given by my teammates, which we need to determine all the methodology, data collections, etc. We also need to fill up this form, so I want you to evaluate and answer me what should the answers to each be as well. For reference we are 4 teammates where Richeek has experience with UI UX Design, SEO, Marketing campaigns, those sort of work, Kushgra is willing to work on llms and other technical stuff, lakshit too on the technical work and also data processing stuff. Finally I think there needs to someone to bring everyone together and manage all the stuff by understanding all and making critical decisions both technical and non-techincal, and I'll be willing to take up this task. So for the final project we'll be buidling a medium fidelity or if possible a high fidelity prototype of the same too. So we need proper actionable and exact decision processs, timeline, exact deliverables, collaborations, and everything we need to effectively complete this project according to the grading rubric too. Therefore get me an indepth report and all the deliverables to what I have asked for and mentioned throught the whole prompt.

Answer

Project Execution Report — BITS F364 Human-AI Interaction Project, First Semester 2026-27

TL;DR

  • Build a claim-level evidence-verification "confidence checker" for enterprise analysts, evaluate it as a controlled within-subjects user study on a business-document/audit task, and make the "actionable advice" layer your central novelty — it is under-explored in published HCI work and testable as a distinct condition, though feasible only at pilot scale (N≈40–60, not N=400).
  • Two of the team lead's motivating anecdotes must be corrected before they enter a graded report: the Subnautica 2 dispute involves KRAFTON (South Korea), not Tencent, with a $250M earnout; the "AI advised firing" claim is a plaintiff allegation (Krafton allegedly asked ChatGPT for ways to avoid the bonus), NOT an established fact. Arthur Andersen/Enron predates AI and is a weak fit — replace it with the Deloitte Australia refund, Mata v. Avianca, the Charlotin hallucination database, and the Lancet fabricated-citation study.
  • The enterprise "AI output verification" market is already crowded with funded incumbents (Patronus AI, Galileo, Arize, Braintrust, Guardrails AI), so a generic hallucination-detection startup has a weak moat — but the human-facing "actionable recourse + trust-calibration" angle, plus GikaGraph.ai as a channel/design partner, is a defensible research and early-commercial niche.

Key Findings

Fact-checks of the two motivating anecdotes (accuracy-critical)

(a) Subnautica 2 — CORRECTED. The publisher is KRAFTON, Inc., a South Korean company — NOT Tencent. Krafton acquired Unknown Worlds Entertainment in 2021 for ~500Mwithanadditionalearnoutofupto∗∗500M with an additional earnout of up to **250 million** tied to revenue milestones. The earnout formula was highly leveraged: per DarrowEverett LLP's legal analysis, "for every dollar above a $69.8 million threshold, Krafton would owe $3.12, up to the cap" — a sum reported by KED Global as roughly 35% of Krafton's 2025 operating profit, which explains the intensity of the dispute. The base testing period ran to December 31, 2025, with seller extension options. In July 2025, Krafton removed the three founders/leaders — Ted Gill (CEO), Charlie Cleveland, and Max McGuire — and delayed the game to 2026. The founders (via Fortis Advisors) sued; Krafton countersued alleging the founders abandoned the project and misappropriated data. On March 16, 2026, the Delaware Court of Chancery (Vice Chancellor Lori W. Will) ruled Krafton breached the Equity Purchase Agreement by terminating the leaders without cause, reinstated Gill as CEO, and "equitably extended the earnout testing period by 258 days, moving the base deadline from December 31, 2025 to September 15, 2026, with the right to extend it further to March 15, 2027." The parties later settled, dismissing all proceedings, expanding the bonus pool to all employees, with Gill agreeing to step down.

The AI claim is an ALLEGATION, not a documented fact. Court filings published by co-founder Charlie Cleveland allege that Krafton — "after announcing their move to become an AI-first company … asked ChatGPT for legal ways to avoid paying the bonus." Per court findings, Krafton CEO Chang Han (CH) Kim allegedly formed a secret "Project X" designed to either negotiate a "Deal" on the earnout or execute a "Take Over" of Unknown Worlds. Krafton publicly rebutted the AI narrative to Eurogamer, calling it "a distraction." What is documented: the allegation exists in Delaware Chancery pleadings, and Krafton's AI-first corporate strategy is real. What is NOT documented: that an AI actually recommended firing the developers, or that AI advice caused the decision. The team must present this precisely as "an allegation in pleadings" and never as established fact. Framed correctly, it is actually a strong motivating case, because it dramatizes exactly the failure mode the project targets — an executive potentially over-relying on an AI's advice on a high-stakes, high-value decision.

(b) Arthur Andersen / Enron — VERIFIED but a WEAK FIT. Facts confirmed: Arthur Andersen was Enron's auditor; Enron collapsed in December 2001 amid off-balance-sheet special-purpose entities; Andersen was convicted of obstruction of justice in June 2002 for shredding documents (fined $500,000, 5 years' probation), which destroyed the firm (from ~28,000 employees to ~200); the U.S. Supreme Court unanimously reversed the conviction on May 31, 2005 (Arthur Andersen LLP v. United States, 544 U.S. 696) on flawed jury instructions. Assessment: this predates AI entirely and is a stretch as a motivating example for an AI-verification tool. Use it only as a brief analogy for "how a due-diligence/verification failure can destroy a firm and harm stakeholders," not as an AI case.

Better, documented AI-harm cases to cite instead

  • Deloitte Australia refund (Oct 2025) — the flagship B2B case. Deloitte partially refunded the Australian Department of Employment and Workplace Relations for an A439,000( US439,000 (~US290,000), 237-page "Targeted Compliance Framework Assurance Review" that contained AI-generated errors — a fabricated quote from a federal court judgment and references to nonexistent academic papers. Deloitte refunded "over 97,000 Australian dollars ($63,000 USD)," the final contract installment. The revised report disclosed use of a "generative AI large language model (Azure OpenAI GPT-4o) based tool chain." Errors were first flagged by Sydney University's Dr. Christopher Rudge: "I instantly knew it was either an AI hallucination or a very well-kept secret." This is the single best case for the B2B framing — a Big Four firm, a government client, real money, fabricated citations, in an assurance/consulting context.
  • Mata v. Avianca (SDNY, June 22, 2023): Judge P. Kevin Castel sanctioned attorneys Steven Schwartz, Peter LoDuca and their firm $5,000 (678 F. Supp. 3d 443) for a brief with six fabricated ChatGPT-generated case citations. The founding example of AI hallucination in professional practice.
  • Court hallucination tracking databases: Legal researcher Damien Charlotin's AI Hallucination Cases database documented 1,598 cases worldwide as of June 9, 2026 (1,668 by July 2, 2026) — "140 new documented cases, just under eight per day" (HAQQ tracker), up from ~200 a year earlier. Record penalty to date: $110,204.38 in Couvrette v. Wisnovsky (D. Or.), across orders in Dec 2025 and Mar 2026, for 15 nonexistent cases and 8 fabricated quotations — more than 20 times the original Mata sanction. This gives the team quantitative evidence of the problem's scale.
  • Academic fabricated citations: Topaz et al., The Lancet (Vol 407, May 9, 2026) audited 2,471,758 papers / 97.1M verified references and found 4,046 fabricated citations across 2,810 papers: "the fabrication rate increased more than 12 times, from approximately four per 10,000 papers in 2023 … one in 2,828 papers in 2023, to one in 458 papers at the end of 2025," and 1 in 277 in the first seven weeks of 2026. Roughly 53 NeurIPS 2025 papers (~1%) carried fabricated citations that survived peer review. Supports the research-lab customer segment.

GikaGraph.ai — stakeholder profile (verified)

  • Company/product: GikaGraph.ai (branded "GiKA AI" / "GiKA") is an "Entity Intelligence Platform" that turns fragmented enterprise data into knowledge-graph-grounded ("Context Graph") insights using specialized small language models, marketed as an "intelligent decision agent" for enterprise decisions with a strong financial-sector focus (investments, portfolio management, market intelligence, customer risk). It pitched at Global Fintech Fest 2025 (Oct 9, 2025, Mumbai) under "Wealth Management" (self-reported via founder's LinkedIn).
  • Founder: Dr. Manoj Agarwal, ex-Senior Staff Engineer at Uber AI (knowledge graphs/semantic search for Uber Eats), Principal Applied Scientist at Microsoft AI & Research (Microsoft Product Knowledge Graph), earlier IBM Research, PhD in data mining/IR. Legal entity Gika AI Private Limited, incorporated June 1, 2024, ROC Bangalore (CIN U62090KA2024PTC189274), registered at Indiqube-Edge, Panathur, Bengaluru 560103; directors Manoj Kumar Agarwal and Rashmi Agarwal. Website also lists a San Jose, CA address.
  • Funding: Unfunded / bootstrapped per Tracxn as of 2025-26; no seed/angel/accelerator round documented; paid-up capital ₹100,000 (founder capital). An advisor, Abhishek Tiwari, joined (per founder's LinkedIn). This matters: an unfunded early startup can realistically offer data, domain expertise, and a letter of support — but NOT seed funding.
  • BITS/APPCAIR relationship: GikaGraph co-organized the JAW 2026 Hackathon ("Quantifying Trust & Reliability in Generative AI Output," 14-15 Aug 2026, BITS Pilani Pilani Campus) with APPCAIR (Anuradha & Prashanth Palakurthi Centre for AI Research, headed by Prof. Snehanshu Saha). The BITS faculty contact is Dhruv Kumar (Assistant Professor, CS&IS; ex-Microsoft Research; PhD, U. Minnesota). Prize pool ₹25,000 (₹15k/₹10k). The team won the finals on 13 Aug 2026. Note: the specific "687-document synthetic construction-company archive" detail could NOT be independently verified — the official page describes only a generic multimodal PDF/XLSX dataset. Confirm the corpus details directly with GikaGraph before citing them.

Alignment with the project concept: Strong thematic alignment (reliability/trust in GenAI output for enterprise decisions; knowledge-graph grounding = provenance) but a different technical layer (GikaGraph builds the reliable-answer engine; the team builds the human-facing trust-calibration/verification interface on top of AI outputs). Complementary, not competitive.

The published methodology base is sound; the key cautionary result is PaperTrail

The team's internal literature review is accurate and well-chosen. The most important finding to design against is Martin-Boyle et al. (PaperTrail, CHI 2026): a claim-evidence provenance interface for scholarly LLM Q&A (within-subjects, N=26 researchers) significantly lowered trust but did NOT change reliance behaviour — participants kept relying on LLM edits to avoid costly verification, and found the interface cluttered under time pressure. Combined with Cao, Liu & Huang (2024) (calibrated frequency framing beats bare confidence numbers) and Kim et al. (CHI 2025) (most interventions reduce over-reliance but fail to improve appropriate reliance; only a reliance disclaimer helped, with alarm-fatigue risk), the evidence says: an interface that merely displays verification signals may change attitudes without changing behaviour. The team's actionable-advice layer is a direct, well-motivated response to this gap.

Details

SECTION 1 — GikaGraph.ai as stakeholder/partner

(i) Alignment. GikaGraph's mission ("prove and trust AI outputs … multi-hop reasoning, numerical consistency over heterogeneous data") is essentially the engine-side of the same problem the team addresses on the human/interface side. Their knowledge-graph grounding produces exactly the kind of provenance the team's claim-level verification UI needs to display. The hackathon problem the team already solved (plain-English question → single number from a document archive) is a numeric-claim-verification task — a natural corpus and task template for the study.

(ii) Realistic partnership forms. Given GikaGraph is an unfunded ~2024 startup that the team just impressed by winning its hackathon, realistic asks are: (1) access to anonymised or synthetic enterprise document corpora (the hackathon dataset is ideal and already familiar); (2) 1-2 domain-expert interviews with client-facing analysts for stakeholder requirements validation (satisfies the rubric's "real affected stakeholders"); (3) a letter of support/collaboration for the CHI SRC and course; (4) recruitment help — a handful of their analysts as expert participants; (5) mentorship from the founder.

(iii) Overreach to avoid: seed funding, equity discussions, exclusive IP arrangements, or large amounts of engineering time. These are premature and inappropriate for a course project with an unfunded partner.

(iv) How to structure the ask given the rubric requires real-stakeholder validation: Frame it as a lightweight, time-boxed collaboration — one 45-minute scoping call to co-define the analyst task and confirm the corpus; a signed data-use/NDA note; two 30-minute expert interviews; and an offer of co-authorship or acknowledgment on any resulting paper. Keep GikaGraph's effort under ~4 hours total. Confirm the "687-document" corpus details and permission to use it, in writing.

Best strategic role for GikaGraph: Partner + channel/design-partner customer, not acquirer or competitor. As an unfunded startup they cannot acquire; they are not a competitor (different layer of the stack); they are an ideal early design partner and source of real stakeholders.

SECTION 2 — Methodology and study design (building on the internal review)

1. Adapting the design to the B2B framing — recommended task and corpus. Recommendation: use a business-document / financial-audit / bid-analysis corpus (the GikaGraph-style archive), NOT a research-literature corpus. Justification: (a) it matches the B2B vision and gives real stakes ("miss a material irregularity → bad business decision"); (b) the team already has domain familiarity and a ready corpus from the hackathon; (c) GikaGraph can supply real stakeholders for validation; (d) it differentiates from PaperTrail (which already covered the scholarly-Q&A setting), strengthening novelty. Concretely: the AI produces an analytical answer (e.g., "Total contingent liabilities across these contracts are ₹X and none breach covenant thresholds"; or "Vendor A's bid is compliant with all mandatory clauses"), and the participant, playing an analyst, must decide whether to accept, revise, or escalate. Ground truth is controllable because the team constructs/curates the corpus and seeds known irregularities.

2. The "actionable advice" mechanism (the team's strongest novelty claim). Prior work exists on algorithmic recourse (Ustun/Wachter-style counterfactuals; personalized recourse; FAccT'24 work on recourse presentation) but it addresses how a subject changes their own attributes to flip a model decision about them — a different problem. There is comparatively little published HCI work on "what to verify next" recourse for a user judging an AI's factual output: i.e., surfacing evidence gaps and recommending concrete verification actions ("Claim 3 is unverified — check the FY24 audit annexure, section 4"; "Two sources conflict on the penalty clause — escalate to legal"). Assessment: (i) Genuinely novel-ish: combining claim-level verification with a ranked, actionable "next verification step / decision recommendation" layer, tested as a distinct condition, is not something the cited literature has isolated and evaluated. This is a legitimate CHI SRC contribution. (ii) Feasible in a semester: yes, because the actionable advice can be generated by templated rules over the verification output (unverified → "retrieve/confirm source X"; conflicting → "reconcile A vs B / escalate"; supported-but-weak → "corroborate with second source") rather than a novel model. (iii) Experimentally testable as a distinct condition: yes — add condition (5) "claim-level + calibrated frequency + actionable advice" to isolate the incremental effect on appropriate reliance and, crucially, on verification behaviour (the DV PaperTrail found unmoved). Core hypothesis: actionable advice converts attitude change into behaviour change (more targeted source click-through, higher incorrect-claim rejection) without adding clutter/workload.

3. Sample size and recruitment.

  • Defensible N for a semester pilot: aim for N≈40-60 in a within-subjects design (each participant sees all/most conditions across counterbalanced trials), which maximizes power per participant. A between-subjects design would need ~31/condition at d=0.3, power .80 (i.e., ~120+ for four conditions) — not feasible; hence within-subjects or a mixed design (interface condition within-subjects; corpus domain or AI-correctness manipulated within-subjects at the trial level). Report it explicitly as an underpowered pilot, pre-register, run an a priori power analysis (R pwr), and report effect sizes with confidence intervals rather than leaning on p-values.
  • Recruitment in India: Prolific is effectively unavailable for recruiting participants in India (historically near-zero Indian participant pool), so do NOT budget for it. Use instead: (a) BITS Pilani MBA / economics / M.Sc. Finance students and senior undergrads as analyst-proxies; (b) chartered-accountancy (CA) articles/students in the region — a strong, motivated, domain-relevant pool; (c) GikaGraph's client-facing analysts for a small expert sub-sample (high ecological validity, satisfies rubric); (d) if remote is needed, Amazon Mechanical Turk India or university mailing lists rather than Prolific. Compensate fairly (e.g., a modest gift voucher).

4. Ethics / IRB / DPDP and the deception condition.

  • What's required at BITS: a student HCI study with human participants should go through the department/institute ethics review (confirm the exact route with the instructor). Require written informed consent, voluntary participation, right to withdraw, data anonymisation, and secure storage.
  • DPDP Act 2023 (Rules notified Nov 2025; core compliance deadline 13 May 2027): collect minimal personal data, give a clear consent notice (purpose, data collected, retention, withdrawal, grievance contact) before collection, keep data pseudonymised (participant IDs, not names), and delete after analysis. Straightforward for a small academic study.
  • The deception (deliberately miscalibrated/fabricated evidence labels) condition: deception is permissible in research when (a) scientifically necessary, (b) risks are minimal, (c) no less-deceptive alternative exists, and (d) participants are debriefed. Here the deception is central to the dark-pattern/persuasion hypothesis (does the interface's mere presence drive trust regardless of content). Recommendation: INCLUDE a mild version, carefully. Do not fabricate evidence about real named entities; use the synthetic corpus so no real party is defamed; keep the miscalibration bounded; run a thorough debrief disclosing the manipulation, explaining why, offering data withdrawal, and checking for lingering misconceptions. Flag it explicitly in the ethics application. If the committee balks or the timeline is tight, fall back to a "presence-only" control (an interface that shows verification chrome but with uninformative/neutral labels) that tests the same "mere presence" effect with less ethical load.

SECTION 3 — Submission-ready answers to the required Google Form

Team Name (options):

  1. TrustLens (recommended — captures "seeing through" AI output)
  2. VerifAI
  3. Recourse
  4. Groundwork (evidence-grounding + doing the groundwork)
  5. BITS & bytes (existing hackathon team identity — good for continuity/branding)

HAI Problem Statement: Knowledge workers in high-stakes enterprise settings (audit, financial analysis, bid/contract review) increasingly rely on LLM-generated analyses, but LLMs produce fluent, confident outputs whose factual grounding is opaque, causing both over-reliance (accepting fabricated or unsupported claims) and under-reliance (dismissing correct output). Existing confidence scores are poorly calibrated and, even when interfaces surface evidence, they change users' stated trust without changing their verification behaviour. There is a need for an interface that not only flags per-claim reliability against real evidence but tells the user what concrete action to take to raise reliability or make a safe decision.

System / Interface Description: A "confidence-checker" web interface that takes an AI-generated analytical answer about a set of enterprise documents, decomposes it into discrete factual claims, verifies each claim against a fixed curated corpus via retrieval + entailment classification (supported / partially supported / conflicting / unverified) with the source passage available on demand, presents calibrated-frequency reliability information, and generates an actionable advice panel recommending the specific next verification step or decision for each weak/conflicting claim. Evaluated against a plain-answer control, an ordinary-citation baseline, and a claim-level-labeling-only condition.

Evaluative Research Question: Does adding an actionable-advice layer (concrete "what to verify / what to decide" recommendations) on top of claim-level evidence verification improve users' appropriate reliance on AI output — and their verification behaviour — compared to evidence labels alone?

Target User Population: Enterprise knowledge workers who make or support decisions using AI-generated analysis — financial/audit analysts, consultants, bid/procurement reviewers, and research analysts — operationalised for the study via domain-adjacent proxies (MBA/finance students, CA students/articles) plus a small sample of practising analysts (via GikaGraph).

Roles and Responsibilities Distribution:

  • Aditya Pradhan (Team lead) — Technical Architecture Lead. Sets overall system design and holds final say on technical and non-technical calls; implements the claim-extraction pipeline and behavioural-data logging inferred from data analysis; coordinates integration across the team's components; solidifies the literature-grounded justification and design-decision reflection sections of the report.
  • Kushagra Agnihotri — LLM/RAG Engineering. Builds the retrieval infrastructure (corpus chunking, embeddings, vector store, hybrid dense + keyword search), the entailment-classification pipeline, the calibrated-frequency computation, and the actionable-advice generation logic — full technical ownership of the verification subsystem end-to-end.
  • Richeek Mishra — UI/UX Design & Front-End Lead. Owns the interface, and leads the UI decisions, built around claim-status color-coding, on-demand source-passage expansion, and a progressive-disclosure layout designed specifically against PaperTrail's clutter problem; runs the internal usability pilot and drives the resulting fixes; owns teaser and demo videos, the poster, and presentation/marketing materials.
  • Lakshit Tanwar — ML/DL for Claim Generation & Analysis. Trains/fine-tunes embedding-similarity models to generate and screen the study's seeded fabricated and conflicting claims; builds the ML-based classifiers used to score behavioural outcomes from logged data; owns participant recruitment logistics and works on quantitative data analysis with Aditya.

Primary Evaluation Method: User Study (controlled, within-subjects/mixed, task-based experiment with behavioural + subjective measures). Justification: the research question is about behaviour and appropriate reliance, which a survey or heuristic evaluation cannot measure; A/B testing lacks the controlled task and ground truth needed; only a task-based user study captures accept/reject decisions, verification actions, and decision accuracy. Supplement with a brief expert heuristic review of the interface during formative design.

IRB / ethics / consent needed? Yes. Human participants, behavioural data, and a deception/miscalibration manipulation all require ethics review, written informed consent, and (for the deception condition) structured debriefing; DPDP-compliant consent notice and data handling apply.

Primary metrics/outcomes: correct-claim acceptance rate; incorrect-claim rejection rate; over-reliance rate; under-reliance rate; composite appropriate-reliance score; final task/decision accuracy; verification behaviour (source click-through, targeted vs untargeted); time per trial; NASA-TLX workload; subjective trust; and post-hoc structured debrief responses (to detect the PaperTrail trust/behaviour dissociation).

Support needed from instructor/TA (Prof. Siddharth Mehrotra): Prof. Mehrotra's own research is precisely appropriate trust in Human-AI interaction (PhD TU Delft under Tielman & Jonker; postdoc U. Amsterdam; work on integrity-based explanations and "inverse trust"; industry experience at Google Research, `Siemens, Mazda; founder of the Safe & Trusted AI group at BITS and TrustAxis Pvt. Ltd.). Specific, strategic asks: (1) methodological review of the appropriate-reliance metric and study design — his core expertise; (2) guidance on the deception condition's ethics and the BITS ethics-approval route; (3) help brokering/validating the GikaGraph stakeholder engagement and possibly a warm introduction; (4) advice on CHI SRC framing and positioning against PaperTrail and the CHI 2025 reliance-intervention literature; (5) a pointer to statistical support for an underpowered within-subjects analysis. Do NOT ask him to write code or recruit participants.

SECTION 4 — Execution plan (mid-Aug to early Dec 2026)

Course checkpoints: comprehensive exam 13 Dec 2026; group report due ~2 weeks before compre (~29 Nov 2026); end-of-ideation report ~1 week before mid-sem exam; end-of-prototyping report ~2 weeks before compre; presentation-day content = teaser video (2-3 min), demo video (2-3 min), scientific poster. Rubric weights to earn: stakeholder alignment + documented requirements (24%), literature/expert-grounded strategy (14%), reflection/justification of design (14%), UX validation incl. validity threats & method description (24%), sustainability/feasibility/technical quality/societal relevance (24%); plus intermediate deliverables (15% of final grade), presentation content (10%), team management (10%).

  • Week 1 (mid-Aug): Kickoff, finalize B2B scope, assign roles, set up repo/Kanban. Deliverable: project charter + roles. Owner: Aditya. Dependency: none. → earns team-management.
  • Week 2: Literature consolidation + GikaGraph scoping call; confirm corpus & data-use permission. Deliverable: requirements list v1 + signed data note. Owner: Aditya (call), Kushagra (lit). Dependency: GikaGraph availability. → stakeholder alignment (24%), literature (14%).
  • Week 3: Stakeholder interviews (1-2 GikaGraph analysts) + draft study design & conditions. Deliverable: requirements list v2, study protocol draft. Owner: Aditya + Lakshit. Dependency: Week 2. → stakeholder alignment.
  • Week 4 (end-of-ideation, ~1 wk before mid-sem): End-of-ideation report — problem, stakeholders, RQ, conditions, metrics. Owner: all, led by Aditya. Dependency: Weeks 2-3. → intermediate deliverable (15%), reflection (14%). GO/NO-GO GATE 1: if GikaGraph corpus not secured, switch to a self-built synthetic audit corpus (fallback locked here).
  • Weeks 5-6: Corpus curation + ground-truth seeding (known irregularities); build claim-decomposition + retrieval pipeline. Deliverable: working extraction + retrieval over corpus. Owner: Lakshit (corpus), Aditya+Kushagra (pipeline). Dependency: corpus. → technical quality (24%).
  • Weeks 7-8: Entailment classifier + calibrated-frequency computation + actionable-advice rule layer; UI v1. Deliverable: end-to-end prototype (all 4-5 conditions togglable). Owner: Kushagra (verification/advice), Richeek (UI). Dependency: Weeks 5-6. → technical quality, design justification. GO/NO-GO GATE 2: if claim-extraction/entailment manual-QA accuracy < ~80% on a 50-claim audit set, fall back to Wizard-of-Oz / pre-scripted claim labels for the study (decouples the experiment from pipeline quality).
  • Week 9: Instrument behavioural logging (accept/reject, click-through, timing); pilot with 3-5 internal users; finalize ethics application + consent + DPDP notice. Deliverable: logged pilot data + submitted ethics app. Owner: Aditya (logging), Richeek (usability fixes), Lakshit (ethics paperwork). Dependency: Week 8. → UX validation (24%), ethics.
  • Week 10 (end-of-prototyping, ~2 wks before compre): End-of-prototyping report + fix clutter/usability issues from pilot. Owner: all. Dependency: Week 9. → intermediate deliverable (15%), validity-threats discussion (24%). GO/NO-GO GATE 3: if ethics approval delayed, run the non-deception "presence-only" control variant.
  • Weeks 11-12: Run main study (N≈40-60); rolling data QA. Deliverable: complete dataset. Owner: Lakshit (recruitment/sessions), Aditya (data). Dependency: ethics approval, prototype. → UX validation (24%).
  • Week 13 (~late Nov): Analysis (appropriate-reliance, verification behaviour, workload); write final report (65%); build teaser + demo videos + poster. Deliverable: final report + presentation assets. Owner: Aditya (analysis/methods), Richeek (videos/poster), all (writing). Dependency: Weeks 11-12. → all report rubric components + presentation (10%).
  • Report due ~29 Nov; presentation day; compre 13 Dec. Buffer + rehearsal in the final days.

SECTION 5 — Technical build spec (standard LLM APIs, no training)

Architecture (pipeline):

  1. Input: AI-generated answer + reference to the fixed corpus (30-100 curated documents).
  2. Claim decomposition: prompt an LLM (e.g., GPT-4o-class or Llama-3.1-70B via API) to split the answer into atomic, self-contained, verifiable claims (decontextualized — resolve pronouns/ellipsis). Follow the FActScore/VeriScore "decompose-then-verify" pattern; extract only verifiable claims (VeriScore-style) to avoid checking opinions.
  3. Retrieval over fixed corpus: chunk documents (~200-400 tokens, overlap); embed with a standard embedding model (OpenAI text-embedding-3, or open bge/e5); store in a lightweight vector DB (FAISS or Chroma locally; pgvector if hosted). For each claim, retrieve top-k passages (k≈5). Add BM25/keyword hybrid for numeric/exact matches (critical for financial figures).
  4. Entailment-style verification: for each (claim, retrieved-passage) set, prompt the LLM to classify supported / partially supported / conflicting / unverified with a required citation to the specific passage and a one-line rationale. Aggregate multiple passages (harmonic-mean or "any-conflict-wins" rule). Consider MiniCheck/NLI as a fast secondary check. Never present token probability or verbalized confidence as "probability the answer is true" (Xiong et al.: LLM verbalized confidence is systematically overconfident).
  5. Calibrated-frequency layer: convert verification results into a frequency framing ("Of 100 answers with this evidence profile, ~N would be fully correct") derived from the team's own held-out labeled set, not from model self-confidence.
  6. Actionable-advice generation: rule-templated over verification labels — unverified → "retrieve/confirm in <doc/section>"; conflicting → "reconcile source A vs B / escalate to <role>"; partially supported → "corroborate with a second source"; plus a decision recommendation (accept / revise / escalate). Rank by claim materiality (seeded severity).
  7. UI: claim list with color-coded labels, on-demand source passage (expand-in-place, not modal overload), calibrated-frequency badge, and a compact actionable-advice panel. Design against PaperTrail's clutter finding: progressive disclosure, default-collapsed evidence, one primary recommended action per weak claim, and a time-pressure-friendly layout. Owned by Richeek.

Controlling ground truth for the experiment: because the team curates the corpus and writes the AI "answers," seed a known mix of correct claims, unsupported/fabricated claims, and conflicting claims with documented gold labels and materiality tags. This makes accept/reject scoring objective.

Behavioural instrumentation: log every claim-level accept/reject toggle, every source expand/click (with claim ID and timestamp), per-trial start/end times, final decision, and condition — to a structured store (e.g., JSON events → CSV/SQLite) keyed by anonymised participant ID. NASA-TLX and trust items via an embedded form. This makes accept-rate, rejection-rate, click-through, and time-on-task fully automatic.

Failure modes & mitigations: (a) claim-decomposition errors → manual QA gate (Gate 2) and Wizard-of-Oz fallback; (b) retrieval misses on numbers → hybrid BM25+dense; (c) entailment over/under-calling "supported" → conservative aggregation, human-audited prompt on a dev set; (d) latency/API flakiness during sessions → precompute all verification results offline before sessions (study answers are fixed, so no live inference needed); (e) API cost → cache aggressively, use smaller models for decomposition.

SECTION 6 — Startup evolution path (critical assessment)

Who would pay: audit/assurance firms (post-Deloitte), consulting firms, financial-services compliance/model-risk teams, pharma/research labs, and legal (post-Mata). The regulatory tailwinds are real: EU AI Act high-risk obligations (baseline 2 Aug 2026, with parts potentially shifting toward Dec 2027 under the Digital Omnibus proposal — the trilogue is active, so treat dates as moving), ISO/IEC 42001:2023, NIST AI RMF (+ July 2024 GenAI profile), and India's DPDP Act (Rules Nov 2025, compliance by 13 May 2027). These push enterprises toward auditable AI-output verification.

Competitive landscape — CROWDED and well-funded. Direct/adjacent incumbents in LLM evaluation, guardrails, hallucination detection and AI observability include Patronus AI (Lynx hallucination model, FinanceBench, regulated-industry focus), Galileo (Luna-2 evaluators, runtime guardrails), Arize (Phoenix observability), Braintrust, Guardrails AI (open-source validators), plus NVIDIA NeMo Guardrails, Arthur AI, and gateway players (Maxim/Bifrost). Honest verdict: the generic "detect hallucinations / score groundedness" layer is already commoditising — the team cannot win there as a course-prototype-turned-startup.

Where the defensible niche is: these incumbents are developer/infra tools that score outputs at runtime; almost none focus on the human-facing decision interface — calibrated trust + actionable recourse that changes an analyst's verification behaviour and decision. That human-factors + workflow-embedded + domain-specific (audit/bid) angle, validated by real HCI evidence, is the genuine (if narrow) moat, reinforced by (a) proprietary domain corpora/benchmarks, (b) integration with a grounding engine like GikaGraph, and (c) published CHI-grade evidence that it actually improves decisions.

Honest gap between prototype and fundable company: large. A course prototype demonstrates an interface and a small-N effect; a fundable company needs enterprise integrations, security/compliance (SOC 2, ISO 42001), a data moat, and evidence of willingness-to-pay. The realistic path is research-first (CHI SRC / workshop paper) → design-partner pilots (via GikaGraph) → decide later whether a standalone product or a feature inside a grounding/observability platform is viable.

GikaGraph as partner/customer/acquirer/competitor: Best as partner and design-partner customer (they supply grounding + real stakeholders; the team supplies the trust-calibration interface layer). Not a competitor (different stack layer). Not an acquirer (unfunded). A future integration or feature-licensing relationship is the most realistic commercial evolution.

Recommendations

  1. Correct the anecdotes immediately. In any graded or published text: (a) attribute the Subnautica 2 dispute to Krafton (South Korea), cite the $250M earnout and the March 2026 reinstatement ruling, and present the AI angle strictly as a plaintiff allegation (Krafton denies it); (b) drop Arthur Andersen as an "AI case" and replace it with Deloitte Australia as the flagship B2B example, backed by Mata v. Avianca, the Charlotin hallucination database, and the Lancet fabricated-citation study.
  2. Lock the task/corpus as the enterprise-audit setting using the GikaGraph corpus (confirm the "687-document" details in writing). This satisfies the "real stakeholders" rubric and differentiates from PaperTrail.
  3. Make the actionable-advice layer the headline contribution and test it as a distinct 5th condition; pre-register the hypothesis that it converts attitude change into verification-behaviour change (the exact gap PaperTrail exposed).
  4. Plan for N≈40-60 within-subjects/mixed, not Prolific. Recruit BITS finance/MBA + CA students as proxies and a few GikaGraph analysts as experts; report honestly as an underpowered pilot with effect sizes and CIs.
  5. Submit ethics early; include a bounded, well-debriefed deception condition on synthetic data, with a non-deception "presence-only" fallback if the committee or timeline requires it.
  6. Use Prof. Mehrotra strategically for methods review, ethics guidance, GikaGraph brokering, and CHI SRC positioning — not for build labour.
  7. Pre-compute all verification results offline so study sessions are fast and API-independent; instrument all behavioural DVs automatically.
  8. Position the startup as research-first and, if pursued, as a human-facing decision layer / GikaGraph integration — not as another hallucination-detection API.

Benchmarks that change the plan: if the GikaGraph corpus/permission isn't secured by end of Week 4 → build a synthetic audit corpus. If pipeline claim-labeling accuracy < ~80% by Week 8 → Wizard-of-Oz labels. If ethics approval slips past Week 10 → run presence-only control. If recruitment stalls below ~30 by Week 12 → collapse to the two most important conditions (evidence-only vs evidence+actionable-advice) to preserve power.

Caveats

  • The "AI advised firing the Subnautica 2 developers" claim is unproven — it is an allegation in Delaware Chancery pleadings; Krafton disputes it. Do not state it as fact.
  • The "687-document synthetic construction-company archive" could not be independently verified; confirm directly with GikaGraph before citing.
  • GikaGraph is an unfunded ~2024 startup — do not assume it can provide funding, large data volumes, or significant engineering time; verify all partnership specifics directly with the founder.
  • EU AI Act dates are in flux (Digital Omnibus / trilogue); treat the Aug 2026 vs Dec 2027 high-risk timeline as provisional.
  • Some competitive-landscape and hallucination-count figures come from vendor blogs and trackers (Charlotin's database, HAQQ, Maxim/Galileo); treat exact numbers as indicative and, where possible, cite the primary database/study (e.g., the Lancet paper, the court orders themselves).
  • The identification of the founder's academic Google Scholar profile is strongly inferred, not officially confirmed (common-name risk).
  • BITS-specific course checkpoint dates are inferred from the assignment description; confirm exact dates with the instructor/handout.
Built with LogoFlowershow