A Compact Screening Protocol for Human-AI Co-creation in Technical Knowledge Work

Authors

  • Alexander D. Gelner AVL Deutschland GmbH, Junkers-Ring 6, 85098 Großmehring, Germany; Technische Hochschule Ingolstadt, Esplanade 10, 85049 Ingolstadt, Germany https://orcid.org/0000-0001-9628-798X
  • Johannes Stahr AVL Deutschland GmbH, Junkers-Ring 6, 85098 Großmehring, Germany
  • Andreas Braun AVL Deutschland GmbH, Junkers-Ring 6, 85098 Großmehring, Germany
  • Alexander Baur Technische Hochschule Ingolstadt, Esplanade 10, 85049 Ingolstadt, Germany https://orcid.org/0009-0003-8194-3691

DOI:

https://doi.org/10.23726/cij.2026.1833

Keywords:

co-creation, human-AI collaboration, knowledge management, LLM assistant

Abstract

Enterprise Large Language Model (LLM) assistants can accelerate access to internal engineering knowledge, but organisations need lightweight ways to decide where such systems should support, be verified, or be kept out of the workflow. We present a compact delegation-screening protocol for human-AI co-creation in document-based engineering work. The protocol combines time-to-answer, category-stratified correctness, and deliberately unanswerable items as an acceptance test for delegation boundaries. We demonstrate a first deployment using 19 internal technical PDF reports, 40 questions spanning eight task types, including four intentionally unanswerable items. Six engineers and a single-shot enterprise LLM assistant answered the same set of questions. The assistant reduced mean time-to-answer from 62.4 s to 9.7 s but achieved lower answerable-item correctness (66.7% vs 78.2%) and answered all unanswerable items confidently instead of abstaining, signalling a governance risk. The method turns these patterns into explicit role-allocation rules for trustworthy human-AI co-creation.

References

Amershi, S., Weld, D., Vorvoreanu, M., Fourney, A., Nushi, B., Collisson, P., Suh, J., Iqbal, S., Bennett, P. N., Inkpen, K., Teevan, J., Kikin-Gil, R., & Horvitz, E. (2019). Guidelines for Human-AI Interaction. In S. Brewster, G. Fitzpatrick, A. Cox, & V. Kostakos (Eds.), Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, pp. 1–13. ACM. https://doi.org/10.1145/3290605.3300233

Becker, D., Deck, L., Feulner, S., Gutheil, N., Schüll, M., Decker, S., Eymann, T., Gimpel, H., Pippow, A., Röglinger, M., & Urbach, N. (2024). Lohnt sich Microsoft 365 Copilot? https://doi.org/10.5281/zenodo.13859937

Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the Dangers of Stochastic Parrots. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, pp. 610–623. ACM. https://doi.org/10.1145/3442188.3445922

Brynjolfsson, E., Li, D., & Raymond, L. (2025). Generative AI at Work. The Quarterly Journal of Economics, 140(2): 889–942. https://doi.org/10.1093/qje/qjae044

Dellermann, D., Calma, A., Lipusch, N., Weber, T., Weigel, S., & Ebel, P. (2019). The Future of Human-AI Collaboration: A Taxonomy of Design Knowledge for Hybrid Intelligence Systems. In T. Bui (Ed.), Proceedings of the Annual Hawaii International Conference on System Sciences, Proceedings of the 52nd Hawaii International Conference on System Sciences. Hawaii International Conference on System Sciences. https://doi.org/10.24251/HICSS.2019.034

Fagadau, I. D., Mariani, L., Micucci, D., & Riganelli, O. (2024). Analyzing Prompt Influence on Automated Method Generation: An Empirical Study with Copilot. In O. Baysal, M. Linares-Vasquez, K. P. Moran, & I. Steinmacher (Eds.), Proceedings of the 32nd IEEE/ACM International Conference on Program Comprehension, pp. 24–34. ACM. https://doi.org/10.1145/3643916.3644409

Geifman, Y., & El-Yaniv, R. (2017). Selective Classification for Deep Neural Networks. https://doi.org/10.48550/arXiv.1705.08500

Gelner, A., Eitel, M., Mikhail, M., Olbrich, L., Pierri, A., Borgato, A., & Landgraf, T. (2023). Exploration on the effectiveness of face-to-face and virtual meetings in educational projects dealing with impact innovation. CERN IdeaSquare Journal of Experimental Innovation, 7(1): 12-17. https://doi.org/10.23726/cij.2023.1415

Huang, L., Yu, W., Ma, W., Zhong, W., Feng, Z., Wang, H., Chen, Q., Peng, W., Feng, X., Qin, B., & Liu, T. (2025). A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions. ACM Transactions on Information Systems, 43(2): 1–55. https://doi.org/10.1145/3703155

Knoth, N., Decker, M., Laupichler, M. C., Pinski, M., Buchholtz, N., Bata, K., & Schultz, B. (2024). Developing a holistic AI literacy assessment matrix – Bridging generic, domain-specific, and ethical competencies. Computers and Education Open, 6: 100177. https://doi.org/10.1016/j.caeo.2024.100177

Kristiansen, J. N., & Ritala, P. (2018). Measuring radical innovation project success: typical metrics don’t work. Journal of Business Strategy, 39(4): 34–41. https://doi.org/10.1108/JBS-09-2017-0137

Lai, V., Carton, S., Bhatnagar, R., Liao, Q. V., Zhang, Y., & Tan, C. (2022). Human-AI Collaboration via Conditional Delegation: A Case Study of Content Moderation. In S. Barbosa, C. Lampe, C. Appert, D. A. Shamma, S. Drucker, J. Williamson, & K. Yatani (Eds.), CHI Conference on Human Factors in Computing Systems, pp. 1–18. ACM. https://doi.org/10.1145/3491102.3501999

Lee, J. D., & See, K. A. (2004). Trust in automation: Designing for appropriate reliance. Human Factors, 46(1): 50–80. https://doi.org/10.1518/hfes.46.1.50_30392

Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W., Rocktäschel, T., Riedel, S., & Kiela, D. (2020). Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. https://doi.org/10.48550/arXiv.2005.11401

Mitchell, M., Wu, S., Zaldivar, A., Barnes, P., Vasserman, L., Hutchinson, B., Spitzer, E., Raji, I. D., & Gebru, T. (2019). Model Cards for Model Reporting. In Proceedings of the Conference on Fairness, Accountability, and Transparency, pp. 220–229. ACM. https://doi.org/10.1145/3287560.3287596

Nambisan, S., Lyytinen, K., Majchrzak, A., & Song, M. (2017). Digital Innovation Management: Reinventing Innovation Management Research in a Digital World. MIS Quarterly, 41(1): 223–238. https://doi.org/10.25300/MISQ/2017/41:1.03

National Institute of Standards and Technology (US). (2024). Artificial intelligence risk management framework. https://doi.org/10.6028/NIST.AI.600-1

Prahalad, C. K., & Ramaswamy, V. (2004). Co‐creating unique value with customers. Strategy & Leadership, 32(3): 4–9. https://doi.org/10.1108/10878570410699249

Raji, I. D., Smart, A., White, R. N., Mitchell, M., Gebru, T., Hutchinson, B., Smith-Loud, J., Theron, D., & Barnes, P. (2020). Closing the AI accountability gap. In M. Hildebrandt, C. Castillo, E. Celis, S. Ruggieri, L. Taylor, & G. Zanfir-Fortuna (Eds.), Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, pp. 33–44. ACM. https://doi.org/10.1145/3351095.3372873

Shneiderman, B. (2020). Human-Centered Artificial Intelligence: Reliable, Safe & Trustworthy. International Journal of Human–Computer Interaction, 36(6): 495–504. https://doi.org/10.1080/10447318.2020.1741118

Additional Files

Published

2026-09-16

How to Cite

Gelner, A. D., Stahr, J., Braun, A., & Baur, A. (2026). A Compact Screening Protocol for Human-AI Co-creation in Technical Knowledge Work. CERN IdeaSquare Journal of Experimental Innovation, 10(2), 135–140. https://doi.org/10.23726/cij.2026.1833

Issue

Section

Part 3: Co-Creating with Machines

Categories