Compare commits
1213 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| 80c1fd0a4c | |||
| 45effb4b7d | |||
| 9303389d84 | |||
| 9514473c23 | |||
| 445ad2273c | |||
| 17f7caeaa3 | |||
| f6ba91b120 | |||
| 65774ce88d | |||
| 13d1815218 | |||
| 41fe113dd0 | |||
| 4970fd6ce4 | |||
| 4bd459e403 | |||
| e37b3d266b | |||
| 94523f946e | |||
| cb0461d177 | |||
| 5e71d09554 | |||
| cb7211b599 | |||
| 43a389d462 | |||
| 656baf8ead | |||
| eca794c0bc | |||
| 7865ec9819 | |||
| f7e8196484 | |||
| bdd228b4db | |||
| 900cab6212 | |||
| 8dc30e4807 | |||
| f9eb3a4f33 | |||
| ec016cde9a | |||
| 564584513f | |||
| a466860fdc | |||
| 03ff10ef6c | |||
| 55b859d05c | |||
| 76f42d8ed7 | |||
| 55e4a850fc | |||
| 5594fb464d | |||
| 64de1cb211 | |||
| 1d3df7fa04 | |||
| ae3263b840 | |||
| e3c99d9964 | |||
| b2bd8218c6 | |||
| 8715a0585a | |||
| 4a95659892 | |||
| 5635fab9d3 | |||
| 92e014a4cc | |||
| 852a8dbad4 | |||
| eada610734 | |||
| 028abc6a74 | |||
| 4c3de2df65 | |||
| cf151f91ff | |||
| 970db97acb | |||
| 4bc942af3c | |||
| 52656a51cb | |||
| 479640e214 | |||
| 650212ef62 | |||
| d7417fb2e7 | |||
| a6a4cc5143 | |||
| 94ba76de0f | |||
| 62560996c2 | |||
| 94fa9e21ba | |||
| 610c2742dc | |||
| 4b8744903a | |||
| fa75ffc404 | |||
| bc3607bae4 | |||
| 9154d836a2 | |||
| 1289fb0bef | |||
| 73bc963398 | |||
| 43e051e725 | |||
| 38bef807f5 | |||
| fc1abe0999 | |||
| dbb3181386 | |||
| 3031b13eaa | |||
| 8f40dde4a4 | |||
| 4bb388c955 | |||
| 29b2acffb1 | |||
| f3d434cdfd | |||
| 33b9a2c9fd | |||
| 272dd18b9f | |||
| cc2998845b | |||
| 035271db60 | |||
| 55eb3dec30 | |||
| f8ba94d416 | |||
| 7aeb39e150 | |||
| 7e9089f270 | |||
| 960e979f74 | |||
| e6a3ed887a | |||
| e86dff2907 | |||
| 12ad112fec | |||
| 250fb97fef | |||
| 62f9416ead | |||
| 76702572ef | |||
| 1b6876a8c2 | |||
| 89f368df03 | |||
| 730ee55e14 | |||
| b4741e077d | |||
| 4a3d33e44f | |||
| 9242902e1b | |||
| b4be17586f | |||
| ec5523d620 | |||
| 1d38492b44 | |||
| f32f613a7f | |||
| de60d05b0c | |||
| cc5a392583 | |||
| 8b04ee0939 | |||
| e19ac4e44d | |||
| c7bcdd4a4a | |||
| 58a89c810f | |||
| 9a0c07f9c7 | |||
| 8619dfda75 | |||
| 683b6e79e5 | |||
| e3746c52d9 | |||
| 0fea7e8347 | |||
| 0a76dd03ce | |||
| f47d486985 | |||
| 1660d306b5 | |||
| ee36d43584 | |||
| a91f630f79 | |||
| bd84b65258 | |||
| 28de3652d3 | |||
| a67d95f58a | |||
| 3a11cf5225 | |||
| 170ee73f94 | |||
| 6f5fbf6dbb | |||
| 516aa0c9d9 | |||
| 0fb2e0944c | |||
| eed9100777 | |||
| f185dfa2a0 | |||
| c5ebf809ea | |||
| 1683357483 | |||
| 0ed4ee6e18 | |||
| 9f361ba7dd | |||
| e8856de1b4 | |||
| 5a10e46f11 | |||
| 8c8a2eb32e | |||
| ff8e3db4b2 | |||
| 8526723b49 | |||
| 0466636b77 | |||
| f903926394 | |||
| 1a1b35d4fb | |||
| fc2d208f30 | |||
| b9cbab149f | |||
| bed924b45d | |||
| e1cb2be4ca | |||
| b1722a7459 | |||
| 0370fd3527 | |||
| 9a2b4a9b85 | |||
| e22f25a431 | |||
| 6e691ee2ea | |||
| ce462354fd | |||
| 02a6b21151 | |||
| 75da8e0200 | |||
| 3d1231e0fe | |||
| 601ecf5503 | |||
| 574a598fae | |||
| 7a5d32bc83 | |||
| 613b8f39a6 | |||
| 9b57f057b4 | |||
| ae60947451 | |||
| 1b7d878b7c | |||
| b80d541946 | |||
| 54ec5f0091 | |||
| 3854c124cb | |||
| 63ebf5ada1 | |||
| fbc5a44045 | |||
| 4bb4400731 | |||
| 4b5a0b89cd | |||
| 9d24382d0e | |||
| 044d44ce0d | |||
| f2fb9ffb66 | |||
| 60b7bee807 | |||
| e9a3e3610c | |||
| ceb238fd1b | |||
| 41c646d898 | |||
| 4b2881c7c4 | |||
| a47b7ea7ec | |||
| 42c3015518 | |||
| ae224b4449 | |||
| 756fa431a7 | |||
| 841f72f296 | |||
| 611d080ff0 | |||
| 48d7e9cae7 | |||
| f7410c8e96 | |||
| da3f15708e | |||
| 498390a741 | |||
| 7833715ac2 | |||
| d996707c0a | |||
| ec99da6375 | |||
| 2d40c09c88 | |||
| 3a3f34f18d | |||
| 7029ea8fff | |||
| 572c7bf1b5 | |||
| 0661c9e9ca | |||
| 05004336a2 | |||
| 8d7f05b6b2 | |||
| ebdb0f2ee1 | |||
| b3688db750 | |||
| 8025ed0b42 | |||
| ba889de480 | |||
| 5df41d3013 | |||
| 2eb8713b53 | |||
| 9a207b6938 | |||
| 9af6ad111c | |||
| 1821bf8094 | |||
| 071e2b68f6 | |||
| c88f339d32 | |||
| 5ffc1ecee4 | |||
| c2cb031461 | |||
| fe3a5e6c27 | |||
| 16e040929b | |||
| 5be06a16a9 | |||
| 81c57c5e7e | |||
| 638388ad17 | |||
| 734d42490a | |||
| 4e43cbaf09 | |||
| fdf2d009a6 | |||
| 333721d72f | |||
| 4c68780ad3 | |||
| 3aad7eba85 | |||
| 4c5112cbf4 | |||
| 106ef05317 | |||
| 2a515f0eb4 | |||
| 9e228fc959 | |||
| bf3e9d178c | |||
| 5d300f06da | |||
| 9d963e8f0c | |||
| 902c5999a3 | |||
| 5515283d5e | |||
| 35c497ebfb | |||
| 66cc689742 | |||
| 12b5471aa3 | |||
| cc57bb1648 | |||
| da82b2cd66 | |||
| cebc76397b | |||
| 82eaf15ad8 | |||
| e80d2d2319 | |||
| 61443ca8af | |||
| 55c8900736 | |||
| 377204714d | |||
| dd3f59e399 | |||
| 67fb85a17f | |||
| f84ef7f649 | |||
| b7ba44688d | |||
| b58d059f4d | |||
| 3ad12bc273 | |||
| 4f3c8a5379 | |||
| 5086b2490e | |||
| b9b0141e27 | |||
| 6b14ebc824 | |||
| 47ecb38d07 | |||
| 48ad4aae19 | |||
| 09ea6aa420 | |||
| 3dffa4ba93 | |||
| 2cef83b477 | |||
| 0a100fb1e4 | |||
| fa049a26fb | |||
| 901348023d | |||
| 83b3833382 | |||
| 5d05d5d5d2 | |||
| bd871cddb2 | |||
| 489221155b | |||
| e3366670ad | |||
| a7ccef4c6a | |||
| 2d665c9a67 | |||
| 86739b1a0f | |||
| 111cbf12ba | |||
| bfc8c6355f | |||
| 690079e2b9 | |||
| fb67680fae | |||
| c06cd45004 | |||
| aeb653e562 | |||
| 0386210b3b | |||
| fa508992ed | |||
| e1d861ca50 | |||
| b0ab25c2d6 | |||
| 7a3f6b70d1 | |||
| 6a89f1b564 | |||
| 76f419f795 | |||
| fe9e70fd4d | |||
| f75a40ebe2 | |||
| c9400550a7 | |||
| 70939dc73d | |||
| 976bccedc4 | |||
| 051c2ea388 | |||
| 40aada1a8b | |||
| 5e0e6d26f0 | |||
| d86c2e27d9 | |||
| 9f5575ada4 | |||
| a7b4851e69 | |||
| 9ed6dadb41 | |||
| ce0d7923b8 | |||
| a4db3e5519 | |||
| 149f5ce631 | |||
| dddbce144e | |||
| c5e132142c | |||
| 7fea1a549e | |||
| c52d25f4f9 | |||
| 570b114770 | |||
| c2f6690ff2 | |||
| 9a96d9e787 | |||
| 67fa4d89a5 | |||
| 8fdb45da69 | |||
| abd5cef032 | |||
| 7dc7fa2be1 | |||
| 13f4a7405e | |||
| 0a5b8c9ce5 | |||
| 881f402aa2 | |||
| 90884968b3 | |||
| 687e974046 | |||
| b8355b2781 | |||
| 1c0ff599c8 | |||
| 8e1915ef60 | |||
| 0cc4805c00 | |||
| 681320eeec | |||
| d00f5424e0 | |||
| 6fc73a4f65 | |||
| 59d74e1c3d | |||
| 58d3ed2303 | |||
| 85929c0b83 | |||
| 14090a2345 | |||
| 8db2ed2d08 | |||
| be9ce37a6c | |||
| 49f593b575 | |||
| 1005106454 | |||
| 34ff8481bb | |||
| f0df572428 | |||
| 72ad8e3391 | |||
| bac71c61da | |||
| a727b1fb40 | |||
| bb817aa905 | |||
| c964926d7c | |||
| aa483fe1df | |||
| 54444ad625 | |||
| 230a9d6759 | |||
| 58374fe0ed | |||
| 47ff1dadcb | |||
| fc609e1e58 | |||
| f25a4f95ed | |||
| 5fa3fd45e4 | |||
| 05dc0685bb | |||
| 1bb5ff6ba5 | |||
| 44bbcfe64b | |||
| 5c9fb2c10c | |||
| c4c1772172 | |||
| 29f422be7a | |||
| f29a6580a1 | |||
| 8f3b10433c | |||
| 22dd2afa28 | |||
| 90009b2793 | |||
| ceb2d5bc00 | |||
| 5e458adbf0 | |||
| 457d1a55c6 | |||
| 2388575f31 | |||
| 910b12df4a | |||
| f61d0c1749 | |||
| 45269bf152 | |||
| a5cd815a34 | |||
| 710449cb2a | |||
| 8c5c6507ce | |||
| f58faa3524 | |||
| 3eeb3b5186 | |||
| 6e790cae41 | |||
| d216493ed1 | |||
| e88bcf2e2b | |||
| 0cf38a63aa | |||
| 9073d5c3f4 | |||
| cae28f063a | |||
| 557b97e062 | |||
| b99d18698b | |||
| 1959dd3ba9 | |||
| e0f6a28c20 | |||
| 615430a87b | |||
| 8ccdea780e | |||
| 158064489e | |||
| 3a5267340a | |||
| a8be7c0056 | |||
| e00f377a7d | |||
| 1f55c42c35 | |||
| eab77a9b0c | |||
| 9ba8787893 | |||
| 1215783363 | |||
| 279a69f056 | |||
| 21da1153ec | |||
| fc6afebc48 | |||
| da1ad76e7e | |||
| 5fd87204e1 | |||
| 4f40c01e35 | |||
| 0f9a37037c | |||
| da124137e6 | |||
| d23d85a353 | |||
| e6f4b363fb | |||
| 2ac2bbb9cb | |||
| 9a1aa7d325 | |||
| c0eac1993c | |||
| 6ac14eb17a | |||
| 8d1094d21b | |||
| bd5a9a92e0 | |||
| 8b755c40d6 | |||
| 7351bf6501 | |||
| a70a002314 | |||
| 696a9a83d8 | |||
| 9ac31ab49f | |||
| 9b1f6eddc8 | |||
| c851f82647 | |||
| 6e05fec62c | |||
| 8cb5fabe4e | |||
| ba3b847ebd | |||
| e32f0f80eb | |||
| a48d80136c | |||
| fc5cc2866a | |||
| 2e1d608bc6 | |||
| 8d3aa15c20 | |||
| d6f950d086 | |||
| cfc54f9d54 | |||
| cda1360e0f | |||
| 148ecabf53 | |||
| bf90fa57dd | |||
| cf441bbafd | |||
| 072410349b | |||
| 253f241847 | |||
| 7fa458e964 | |||
| 933a3a93d7 | |||
| 93da933776 | |||
| c4a87ab887 | |||
| dde45fc9e6 | |||
| 4902b075c4 | |||
| 6b90ab3ac6 | |||
| 81a1c1b5c8 | |||
| c7386ec9c4 | |||
| 2b32271df9 | |||
| bb149733ac | |||
| 43b63dcd8f | |||
| 5980665553 | |||
| ed087897cd | |||
| ef346440a5 | |||
| 122165aea6 | |||
| da3d45c390 | |||
| 6c711b0082 | |||
| a5cf08ca0c | |||
| 74832045c5 | |||
| ee674b8eca | |||
| 02bf923d72 | |||
| 5edf91a6c6 | |||
| d133de18bc | |||
| 7e71ba8902 | |||
| fc284d43e7 | |||
| 5a0eb2d404 | |||
| 78f20f45c7 | |||
| 6865e2172d | |||
| 159beb5613 | |||
| f04ab55194 | |||
| b7cdd78021 | |||
| 6814a54711 | |||
| 4df07f08e5 | |||
| b4d2850f7f | |||
| 45813070c4 | |||
| 0b396daeed | |||
| 09e8f8010d | |||
| ac8fc9d731 | |||
| e550663e7b | |||
| c4a2fcb750 | |||
| 02e8b03df5 | |||
| fabba01525 | |||
| 176dde9085 | |||
| 2fee7ede4b | |||
| 3cefce90a1 | |||
| f2f0b8c21a | |||
| 74e8c1e22d | |||
| 3589b6b7ff | |||
| b96b0ce39d | |||
| 6375b974c6 | |||
| a23fe13246 | |||
| 35aaaf6ed4 | |||
| db4a462acc | |||
| 335b7e64fd | |||
| e8c749e968 | |||
| a226a0d0e9 | |||
| 1be977b867 | |||
| 48cdae42b1 | |||
| 84fa471e5f | |||
| c6f90cf445 | |||
| 1526c28ac6 | |||
| 87d1023eba | |||
| 7fd72fdafb | |||
| 980ca3f6a3 | |||
| e7672649c0 | |||
| cc2ace522d | |||
| 0110b39bd0 | |||
| b3c6700e23 | |||
| 069f126490 | |||
| 5510bda2c3 | |||
| 1cea24114f | |||
| 0c5698b986 | |||
| b6b291f684 | |||
| c473b0ca30 | |||
| bc9dbd7de5 | |||
| 67687cbab3 | |||
| 947fff8bd0 | |||
| 0793a8ce91 | |||
| f76ec8ac1c | |||
| 1168ee1ad1 | |||
| 85e56167f7 | |||
| f1f88f8fa4 | |||
| 72204371c3 | |||
| f1d8923e48 | |||
| 86f5c71e19 | |||
| 2531843df8 | |||
| 9e812de8ab | |||
| 627a160483 | |||
| 35a21e847e | |||
| 877a7c3f4e | |||
| 3a5e6fe4b8 | |||
| c96de5e91e | |||
| 195b53d45b | |||
| a86b8acbd0 | |||
| 7f59aa58cb | |||
| d2e9c9328b | |||
| a53d6bd80d | |||
| 5a43ea7d8e | |||
| 3fc7b09644 | |||
| 7722970661 | |||
| ff3a186ec3 | |||
| 78e6a04ad8 | |||
| 5949f09444 | |||
| 56ec673e91 | |||
| f923b13707 | |||
| 22c57b1341 | |||
| 01b18d028f | |||
| b39ae9c721 | |||
| f3284ba722 | |||
| 9f0b4bb16f | |||
| 880e4c4b75 | |||
| f03b4c86c8 | |||
| 64e411f03c | |||
| 6601c1eca8 | |||
| 48e28a3405 | |||
| c97c95817d | |||
| 4b63ed7b5b | |||
| 01f9c3775e | |||
| 210585e3c3 | |||
| 1fee0d9660 | |||
| a750053576 | |||
| e3c3c34c3a | |||
| 9fd5fcf5de | |||
| 6491c63d72 | |||
| c62981e32f | |||
| a9485c800f | |||
| 17f4843b95 | |||
| d6c4390134 | |||
| 3391ec656a | |||
| 3a20134d2c | |||
| e5ea4d15b8 | |||
| c1432cab34 | |||
| 32c3040195 | |||
| ba44e4a973 | |||
| 9a1543a238 | |||
| 1f2bec7f13 | |||
| 669a3555ce | |||
| d83f093346 | |||
| d463a97fee | |||
| d66d3915cf | |||
| e0bde88e28 | |||
| 9de42f3bf6 | |||
| 1bc96cda27 | |||
| 21ac1d28fd | |||
| f0d3672f66 | |||
| 3ce7c24d34 | |||
| edfdd57321 | |||
| 2b3e693618 | |||
| 78cca95e7f | |||
| 6577387a5c | |||
| 00bc85e8d7 | |||
| 9dc709b707 | |||
| 01b01cbee9 | |||
| cecce0adf6 | |||
| 8490023d15 | |||
| 7a89ad5296 | |||
| 15cab802b3 | |||
| 3e465ed769 | |||
| b497b3d91e | |||
| fe969312d7 | |||
| dfc1f33fda | |||
| 3acfa826be | |||
| 038c84e0bc | |||
| 5ce91f2964 | |||
| fa4849cc93 | |||
| c586f67e1a | |||
| 7875044dbf | |||
| bd33a5258d | |||
| 34ba1d0293 | |||
| 798f041b94 | |||
| f1830f9c1c | |||
| bcb9a08fdd | |||
| 494b474819 | |||
| 63474edf36 | |||
| 7b6a0711ef | |||
| 6ab83d4382 | |||
| 06605826b3 | |||
| 675a09105c | |||
| 587aa17d7b | |||
| d6b1988040 | |||
| 819b716bf2 | |||
| 5959177950 | |||
| 512b9b6a46 | |||
| cdc0ad6b26 | |||
| 88cda62fe6 | |||
| 42f2c401c6 | |||
| e449a2dc7b | |||
| 014dea2c58 | |||
| 8455af0a09 | |||
| 11ec33a553 | |||
| bb727b19c5 | |||
| e833518b8b | |||
| 2abdfc4d12 | |||
| a022a31923 | |||
| 429f365b07 | |||
| 05424cebc6 | |||
| 38560ac26a | |||
| 07a50ee8ab | |||
| 0c80bd5df0 | |||
| 4ef751c949 | |||
| 4e51cc8727 | |||
| 4dd1616874 | |||
| 0b88565cff | |||
| 5c91b4e50e | |||
| 8660857f6b | |||
| c3ccecd3c4 | |||
| 86acfb4f7f | |||
| bf2625cfed | |||
| 4330b2cf1c | |||
| a735aa0b24 | |||
| c980f5bd4e | |||
| fda7f308bd | |||
| b1ffad0efe | |||
| 813b0351d6 | |||
| 5d48424f91 | |||
| 97e811b0f5 | |||
| 729671cf6e | |||
| 2b9f8d164f | |||
| d79bd9823b | |||
| f3f69c21a5 | |||
| dd0de62b1b | |||
| e0b243deac | |||
| c06a8bb42a | |||
| a4bd62a072 | |||
| f8153b95a2 | |||
| da4acc659e | |||
| 9cc895818d | |||
| 87af6e3256 | |||
| ec257bb406 | |||
| 8bc42640f8 | |||
| a459e367e3 | |||
| d4075b6c5d | |||
| 40952505df | |||
| 19b7c0bbab | |||
| a17fc3f752 | |||
| 09443fd044 | |||
| b5d9152a27 | |||
| b4c80b8870 | |||
| 24d276751b | |||
| 78a683ca80 | |||
| ace9f11ec9 | |||
| c35c0afe90 | |||
| 09948333b7 | |||
| 718a99e1e0 | |||
| 02882efad7 | |||
| fd65d1a14d | |||
| 6b58e78751 | |||
| bc374a0c1d | |||
| afa224ba14 | |||
| c4ed605d33 | |||
| 007a65c122 | |||
| ba55bbd596 | |||
| 7404531696 | |||
| 659f706d56 | |||
| a22027dff3 | |||
| 21b6a13830 | |||
| 10b73ae96c | |||
| 93740fa49c | |||
| 07be768fa1 | |||
| 650a85d0f4 | |||
| 15bae27e25 | |||
| 56680f43d6 | |||
| 35236676bf | |||
| 0d695a2f90 | |||
| 0dccd01ed1 | |||
| d9ce628e8b | |||
| 108e148dd6 | |||
| 6f2bcb9a68 | |||
| 714ff330ae | |||
| 0ac51e2a3b | |||
| f7c8c27b82 | |||
| 4b3986b304 | |||
| 71fa12cd87 | |||
| 8e1fd7a5d2 | |||
| 5f2e83ec35 | |||
| 4d7129055a | |||
| 9f96338350 | |||
| df020d130b | |||
| 9fcc68f294 | |||
| e0e6794d36 | |||
| 9edc53faf3 | |||
| d258d8d18b | |||
| a9d95b4fa8 | |||
| 509ddda778 | |||
| a05af4b3ba | |||
| fff8d8905b | |||
| 0d2d77152a | |||
| 7a14b4b65b | |||
| 1173bda294 | |||
| 7b752dcf62 | |||
| db68d1c30e | |||
| a7c539f173 | |||
| 0dae2f0f1f | |||
| 9787117824 | |||
| d4e3dfbdea | |||
| b3074500c5 | |||
| e93d778df9 | |||
| bc5aca3b9c | |||
| c6569cbd89 | |||
| 9bcf0817ef | |||
| e3fe4f4abb | |||
| 283ead8fe8 | |||
| a9e63a31a2 | |||
| 48164ecac2 | |||
| 95d71fa54f | |||
| 2b3bfb80df | |||
| a776d809d3 | |||
| 05dc79da72 | |||
| 2c0f7c25aa | |||
| 6a5d9ce4be | |||
| 767cc00c9f | |||
| 3ac64f7a64 | |||
| 73e7843075 | |||
| aa524e9800 | |||
| d5151338dd | |||
| 69c3357e8f | |||
| 92fcf51abb | |||
| 763048f905 | |||
| 465b45e96f | |||
| 15ad038ffb | |||
| 10514ca913 | |||
| d659a738da | |||
| 0c4f2b9d1f | |||
| 3b9368db56 | |||
| f91b38f9e9 | |||
| be9441998f | |||
| 1466ddbc6c | |||
| 4ba2e8af4f | |||
| 8f4c2cdd34 | |||
| a30c32e722 | |||
| 59c352296d | |||
| 8c4d4d0c25 | |||
| a714cf48c3 | |||
| 4bc33e32f6 | |||
| 986eb7d99b | |||
| 03dca6832c | |||
| e1f267588f | |||
| fdcf6d3416 | |||
| 351104f4d6 | |||
| 648d14dcc1 | |||
| 59a8b0fa51 | |||
| 04eec50397 | |||
| bacc65b143 | |||
| 12405624f3 | |||
| c422030714 | |||
| 9db9c01b91 | |||
| e47a14ff73 | |||
| d1abf43dfe | |||
| 263048967d | |||
| 00ec7127f4 | |||
| 53ec9d55d0 | |||
| cf1b933660 | |||
| 8dfac2aeec | |||
| 368734f4f7 | |||
| 748ac80da7 | |||
| f78d0b5deb | |||
| da32045a91 | |||
| 886b9fcbbf | |||
| 630cd3763c | |||
| d9f1d5f7f8 | |||
| 1a54ce7920 | |||
| 2ccdbdfd3d | |||
| 3b7aed7e7c | |||
| d9a4144ca2 | |||
| 680554cc92 | |||
| cfddc7c214 | |||
| 75fd7918b8 | |||
| f7711afd45 | |||
| 33f3401b0c | |||
| c96a40ac0c | |||
| 1b49f0e1f7 | |||
| 1f705d5efd | |||
| 2efaf4a6aa | |||
| 559a83285e | |||
| 94077432b1 | |||
| 0f9c9e359b | |||
| 9078e29c0c | |||
| 0442b82c27 | |||
| bca9737f56 | |||
| 095496e6ba | |||
| d7de4702f1 | |||
| f3cac17305 | |||
| d86886c983 | |||
| 6816cc78e0 | |||
| 02ebf0ee87 | |||
| a450c21bfe | |||
| 2c1e35a610 | |||
| 85d119bc61 | |||
| 7e007a1efd | |||
| 14d7d591e2 | |||
| d186ded090 | |||
| a5fc25971b | |||
| c76fc269bd | |||
| f80772f675 | |||
| 57a39ae2a2 | |||
| 6fc618e9e3 | |||
| 22a63aad8b | |||
| c02d8638c1 | |||
| e1076f53ea | |||
| 8994c78e95 | |||
| c2ee7c890c | |||
| ac64c02745 | |||
| 7b2ba37866 | |||
| 503a6ea7a5 | |||
| 46d0d2fde7 | |||
| 4f5487806e | |||
| e87552a0d1 | |||
| 4c4b7c2a79 | |||
| 5e1db14da5 | |||
| 4007cbacaf | |||
| a425859c6f | |||
| d91e39cd3a | |||
| e47b47af1b | |||
| 566b188fb5 | |||
| 7a4a22f0a9 | |||
| a4c125ef24 | |||
| 10ceb8bfc8 | |||
| 827af41d04 | |||
| aed6359047 | |||
| e486b3a6bd | |||
| 76d637d949 | |||
| bc949c37a0 | |||
| 85d7d5d1be | |||
| ee751cb203 | |||
| aeaf83f70f | |||
| 605611c542 | |||
| 456d284b4b | |||
| 31ed0913e1 | |||
| 490177059b | |||
| d1465516ac | |||
| c583dfc104 | |||
| efa88f79e7 | |||
| 581111c891 | |||
| 789575f27f | |||
| 5e6269c7af | |||
| 7c497c9ea8 | |||
| cf11d33679 | |||
| ee3919085c | |||
| 796aa2da03 | |||
| e291b10790 | |||
| 625083aab4 | |||
| 916834823a | |||
| f3339f026c | |||
| 6e5010039a | |||
| daa969585a | |||
| 3556d7bcaf | |||
| 2395b681fb | |||
| 5a9cab876b | |||
| 9ad2949661 | |||
| f37f70b122 | |||
| b6bec026f5 | |||
| 9e4f992a12 | |||
| 5ef3020452 | |||
| 29d274dc87 | |||
| 50d0ffe8f7 | |||
| dd0b65c387 | |||
| 83bb4d8094 | |||
| 329806a5a8 | |||
| 50a909a373 | |||
| d6f385cef2 | |||
| 0a7bb1b5b5 | |||
| f612516216 | |||
| 9ad148b35c | |||
| b302974d47 | |||
| 053dbc4f08 | |||
| 6df6fe16ee | |||
| 244ce39e06 | |||
| 184757b7b6 | |||
| 8a73b69e1c | |||
| 824a431229 | |||
| 9e01cf749d | |||
| f76bf33745 | |||
| a9bdf8ee1b | |||
| 6904dcbbdb | |||
| 5de3b5849e | |||
| 89dc1fba67 | |||
| cb316772f1 | |||
| 5f560be6c7 | |||
| de9d1fd0c4 | |||
| 6e154bac0d | |||
| 164acb59d9 | |||
| d7eeaf2f60 | |||
| 3b36fd9552 | |||
| 711a2e76e1 | |||
| b5bf795b6c | |||
| a81601a536 | |||
| 0bcf198581 | |||
| 6502a4c52c | |||
| 646c61816a | |||
| 86e26e975f | |||
| e8d311bf9c | |||
| 7dda9d8eca | |||
| e4b3150d28 | |||
| 714ee0d4fe | |||
| ff6d55f03a | |||
| 11b1c1e332 | |||
| 0ab3765220 | |||
| 097b12fb34 | |||
| e382ec0bb8 | |||
| 0a246713d2 | |||
| 36fa8f3e63 | |||
| 71be6789ec | |||
| cca93a0c54 | |||
| 2c5041d343 | |||
| dcf9cf7e60 | |||
| 2b8c408048 | |||
| 8e0366bad8 | |||
| d4e1b60f22 | |||
| df95141c2d | |||
| af2fb63fa2 | |||
| f903ad0ac4 | |||
| 594efb0a06 | |||
| ebc5443b0b | |||
| 7222a84ef6 | |||
| 762447a10c | |||
| a4f7204a69 | |||
| 31bd3acf33 | |||
| a3128cefa9 | |||
| cb72933fdd | |||
| 857c70ee0a | |||
| 107ef043fd | |||
| 16f0f58b3e | |||
| d36beaeb98 | |||
| 992faf730c | |||
| fcb9b50c9e | |||
| 07345d07cc | |||
| c913454289 | |||
| ee12d4294c | |||
| 90a57d0965 | |||
| 7560cabb46 | |||
| a6378ce05a | |||
| 4b229d1001 | |||
| cec3a9af94 | |||
| a17625e220 | |||
| c436389f5e | |||
| a9b8ab3924 | |||
| 184eafcbc7 | |||
| 043b3d61dd | |||
| ab3b85f3ca | |||
| aea05a6075 | |||
| 42a98f5699 | |||
| e976aff8d0 | |||
| 65cb91ce26 | |||
| e4699c358b | |||
| 4c436b7bde | |||
| e3b4856a20 | |||
| 534e2d5ea9 | |||
| bad88e7407 | |||
| ce673115f5 | |||
| f5fcd30527 | |||
| 417c19b38c | |||
| eb5483ba6a | |||
| 975cdece98 | |||
| cb0125c18e | |||
| 5b758bd81a | |||
| 0df79034cf | |||
| 9fe9a5ed67 | |||
| efee1a7268 | |||
| 6c816c3f11 | |||
| 5539afc8c5 | |||
| 27de49a63a | |||
| bc71e8bc01 | |||
| 88b8199754 | |||
| 54e4087d7b | |||
| 237d6724c1 | |||
| 907533a6a1 | |||
| e1971878da | |||
| 5c73311f28 | |||
| 818344c0cd | |||
| 25c61e8c0b | |||
| af76c36c12 | |||
| 1f75464afd | |||
| fb68e77f6b | |||
| 5e31be9aac | |||
| 63d72d041b | |||
| 61e4fb8d51 | |||
| a9763dcbfa | |||
| ab6657877b | |||
| 32fe6ccb5d | |||
| d5205352cf | |||
| f92b561130 | |||
| 473e0a22f2 | |||
| 12102e61b5 | |||
| 3a16a24dac | |||
| a1ad51824e | |||
| cfed1860ee | |||
| 0233cc0e0f | |||
| ef812c5eda | |||
| 6a19e6e829 | |||
| d338f03274 | |||
| 7cc0c31cb0 | |||
| 8541956b15 | |||
| bdc91df9ef | |||
| c37c0076dd | |||
| b459cc4782 | |||
| 114b320078 | |||
| d968d7f62c | |||
| 4cb07d541d | |||
| 078eef274b | |||
| aba88c76cd | |||
| 2224d45f42 | |||
| bd527ad579 | |||
| 656ee0cf2c | |||
| c4d876313c | |||
| 7a4c024989 | |||
| c13b163fec | |||
| 75065bac07 | |||
| a64208c6e6 | |||
| e2d2938b3b | |||
| 8f1ed6c095 | |||
| 7e632c2120 | |||
| 2c89461a47 | |||
| 430dbaa2ac | |||
| 64b4a5f888 | |||
| 00412a15fd | |||
| 7b84678657 | |||
| d052c1f025 | |||
| 44042e165a | |||
| 88818fdace | |||
| a16da90a19 | |||
| 46101cb5ac | |||
| b09d37d832 | |||
| 4350c3440f | |||
| f6a6c51230 | |||
| e0b1cdfc1f | |||
| 4d1ce04482 | |||
| c5d50109c4 | |||
| 68b0ef157c | |||
| 6cb9639cd2 | |||
| 3e100f54b2 | |||
| f20aa40f56 | |||
| 7a8866f83a | |||
| de8accc36d | |||
| 4a31bb6109 | |||
| 4b8f40ecea | |||
| 90465776fc | |||
| 23404e8ff2 | |||
| bc32b99b01 | |||
| fdbc618b2c | |||
| 72bd289f60 | |||
| 362144eb8f | |||
| a25dac771b | |||
| 2f8ea0ae30 | |||
| ae4ba3cb1f | |||
| aeb4381aab | |||
| bbbcdae612 | |||
| 32997b2c73 | |||
| f63cc0cf65 | |||
| 076eb72ff9 | |||
| 00c87ea6a8 | |||
| 99ba26077e | |||
| 66e5be38f4 | |||
| dff0548d13 | |||
| 479c171603 | |||
| 6293d663b6 | |||
| c7d50cb1f1 | |||
| 8fdbe09387 | |||
| 2908eb751a | |||
| 329a1bf0b5 | |||
| 72d0d75348 | |||
| 2da6194e1f | |||
| d6ef221252 | |||
| 38108ea477 | |||
| 14154b78b6 | |||
| 1f2dc816fa | |||
| 5888cf2b9b | |||
| b62016654d | |||
| af4cef3162 | |||
| bd988dc14e | |||
| 3d66226906 | |||
| bbcf039d6d | |||
| e651a298d7 | |||
| befe19db27 | |||
| 741da67b63 | |||
| a4499d4b72 | |||
| eafd8df998 | |||
| aab7bdd30b | |||
| 7a00eddd4a | |||
| c0974d7c0e | |||
| 3a64c6febf | |||
| c078c69415 | |||
| 38c5235c12 | |||
| 3e0122c73e | |||
| 8b7d862454 | |||
| 9d32577168 | |||
| 78675ff94d | |||
| 9222bee26e | |||
| e5d3ad7b99 | |||
| a9757de767 | |||
| 6b94ad0f85 | |||
| db85a6d031 | |||
| c92fe77e83 | |||
| e251fa6b1b | |||
| 1720e4a8bc | |||
| 6358cd9525 | |||
| e5ba91ba04 | |||
| d1d2ecca53 | |||
| f09874c6be | |||
| 586d4b2829 | |||
| 18cb55e4c4 | |||
| 86642716e7 | |||
| fcafae1d50 | |||
| 27912d4af8 | |||
| 017ad69de1 | |||
| 381e90d9b4 | |||
| c772205a7e | |||
| 25db658f75 | |||
| 445033946b | |||
| 0475450d59 | |||
| 816b7702bc | |||
| 271a1a4b97 | |||
| a6580a5090 | |||
| c1cb7ce538 | |||
| 653abc5f59 | |||
| 94f80353d9 | |||
| c2005f8273 | |||
| 7665dd1acc | |||
| 421693ab84 | |||
| 3e5f3c03e5 | |||
| d88bf14360 | |||
| 0f21c8aedb | |||
| c1466c6216 | |||
| 91c62c1dea | |||
| 0cf503e1c2 | |||
| 901d2ac57c | |||
| 2b9b8f7e73 | |||
| 281a7b2bb6 | |||
| 6d2806f10c | |||
| 0eee6b8305 | |||
| 8dfd6ff35c | |||
| dcb88e69cd | |||
| c98e234845 | |||
| 5c7c678aef | |||
| 05db7a68cd | |||
| 4a529e6c16 | |||
| 204bec1b8b | |||
| 4046fcb3fa | |||
| 995af4d83e | |||
| d4c7a23e1d | |||
| 775d3e237e | |||
| 3e7b2863e9 | |||
| cfe9099f3f | |||
| 64383502dd | |||
| ad80f788b9 | |||
| 0a28d71cca | |||
| 16fb29c1f0 | |||
| b699d9a533 | |||
| 47fa8e87b1 | |||
| 71968625cc | |||
| d46e2ec35b | |||
| 6e078bf7a9 | |||
| a96108e279 | |||
| db462e32a3 | |||
| 1364f4408e | |||
| 1992be3e8d | |||
| 079764f0ab | |||
| 9fa5c39d69 | |||
| ce2e2a4571 | |||
| 466b44df18 | |||
| 428c9a65bf | |||
| 003cbfe5f5 | |||
| c282324d95 | |||
| 5fe096df67 | |||
| 1847008e0f | |||
| 02b6e7013c | |||
| 1994f9d4c4 | |||
| 2c46dae377 | |||
| f9763495b8 | |||
| db0ee9d5a5 | |||
| 4187fba955 | |||
| aa197e1e14 | |||
| 8fd7773a5e | |||
| 3bbc7c48cb | |||
| 45eb41f1e6 | |||
| e11b822d5f | |||
| af80e3a971 | |||
| 3755ea8658 | |||
| a113fea0ee | |||
| 178020ea33 | |||
| 111fc9ee66 | |||
| 83ce49ec7e | |||
| 2bdf9b7b11 | |||
| 942ba9840b | |||
| a0254b0b74 | |||
| 0a3dfa071a | |||
| 616d8e7f4b | |||
| e3698f32b1 | |||
| 4b8472da7e | |||
| 5639606163 | |||
| 472e8c13bd | |||
| bd404e0f87 | |||
| 0faadf7f7b | |||
| 65cae71b14 | |||
| 80de53e879 | |||
| ce1abe6006 |
@@ -0,0 +1,38 @@
|
||||
---
|
||||
name: code-change-verification
|
||||
description: Run the mandatory verification stack when changes affect runtime code, tests, or build/test behavior in the OpenAI Agents Python repository.
|
||||
---
|
||||
|
||||
# Code Change Verification
|
||||
|
||||
## Overview
|
||||
|
||||
Ensure work is only marked complete after formatting, linting, type checking, and tests pass. Use this skill when changes affect runtime code, tests, or build/test configuration. You can skip it for docs-only or repository metadata unless a user asks for the full stack.
|
||||
|
||||
## Quick start
|
||||
|
||||
1. Keep this skill at `./.agents/skills/code-change-verification` so it loads automatically for the repository.
|
||||
2. macOS/Linux: `bash .agents/skills/code-change-verification/scripts/run.sh`.
|
||||
3. Windows: `powershell -ExecutionPolicy Bypass -File .agents/skills/code-change-verification/scripts/run.ps1`.
|
||||
4. The scripts run `make format` first, then run `make lint`, `make typecheck`, and `make tests` in parallel with fail-fast semantics.
|
||||
5. While the parallel steps are still running, the scripts emit periodic heartbeat updates so you can tell that work is still in progress.
|
||||
6. If any command fails, fix the issue, rerun the script, and report the failing output.
|
||||
7. Confirm completion only when all commands succeed with no remaining issues.
|
||||
|
||||
## Manual workflow
|
||||
|
||||
- If dependencies are not installed or have changed, run `make sync` first to install dev requirements via `uv`.
|
||||
- Run from the repository root with `make format` first, then `make lint`, `make typecheck`, and `make tests`.
|
||||
- Do not skip steps; stop and fix issues immediately when a command fails.
|
||||
- If you run the steps manually, you may parallelize `make lint`, `make typecheck`, and `make tests` after `make format` completes, but you must stop the remaining steps as soon as one fails.
|
||||
- Re-run the full stack after applying fixes so the commands execute in the required order.
|
||||
|
||||
## Resources
|
||||
|
||||
### scripts/run.sh
|
||||
|
||||
- Executes `make format` first, then runs `make lint`, `make typecheck`, and `make tests` in parallel with fail-fast semantics from the repository root. It also emits periodic heartbeat updates while the parallel steps are still running. Prefer this entry point to preserve the required ordering while reducing total runtime.
|
||||
|
||||
### scripts/run.ps1
|
||||
|
||||
- Windows-friendly wrapper that runs the same sequence with `make format` first and the remaining steps in parallel with fail-fast semantics, plus periodic heartbeat updates while work is still running. Use from PowerShell with execution policy bypass if required by your environment.
|
||||
@@ -0,0 +1,4 @@
|
||||
interface:
|
||||
display_name: "Code Change Verification"
|
||||
short_description: "Run the required local verification stack"
|
||||
default_prompt: "Use $code-change-verification to run the required local verification stack and report any failures."
|
||||
@@ -0,0 +1,208 @@
|
||||
Set-StrictMode -Version Latest
|
||||
$ErrorActionPreference = "Stop"
|
||||
|
||||
$scriptDir = Split-Path -Parent $MyInvocation.MyCommand.Definition
|
||||
$repoRoot = $null
|
||||
|
||||
try {
|
||||
$repoRoot = (& git -C $scriptDir rev-parse --show-toplevel 2>$null)
|
||||
} catch {
|
||||
$repoRoot = $null
|
||||
}
|
||||
|
||||
if (-not $repoRoot) {
|
||||
$repoRoot = (Resolve-Path (Join-Path $scriptDir "..\\..\\..\\..")).Path
|
||||
} else {
|
||||
$repoRoot = ([string]$repoRoot).Trim()
|
||||
}
|
||||
|
||||
Set-Location $repoRoot
|
||||
|
||||
$logDir = Join-Path ([System.IO.Path]::GetTempPath()) ("code-change-verification-" + [System.Guid]::NewGuid().ToString("N"))
|
||||
New-Item -ItemType Directory -Path $logDir | Out-Null
|
||||
|
||||
$steps = New-Object System.Collections.Generic.List[object]
|
||||
$heartbeatIntervalSeconds = 10
|
||||
if ($env:CODE_CHANGE_VERIFICATION_HEARTBEAT_SECONDS) {
|
||||
$heartbeatIntervalSeconds = [int]$env:CODE_CHANGE_VERIFICATION_HEARTBEAT_SECONDS
|
||||
}
|
||||
|
||||
function Resolve-MakeInvocation {
|
||||
$command = Get-Command make -ErrorAction Stop
|
||||
|
||||
while ($command.CommandType -eq [System.Management.Automation.CommandTypes]::Alias) {
|
||||
$command = $command.ResolvedCommand
|
||||
}
|
||||
|
||||
if ($command.CommandType -in @(
|
||||
[System.Management.Automation.CommandTypes]::Application,
|
||||
[System.Management.Automation.CommandTypes]::ExternalScript
|
||||
)) {
|
||||
$commandPath = if ($command.Path) { $command.Path } else { $command.Source }
|
||||
return [PSCustomObject]@{
|
||||
FilePath = $commandPath
|
||||
ArgumentList = @()
|
||||
}
|
||||
}
|
||||
|
||||
if ($command.CommandType -eq [System.Management.Automation.CommandTypes]::Function) {
|
||||
$shellPath = (Get-Process -Id $PID).Path
|
||||
if (-not $shellPath) {
|
||||
throw "Unable to resolve the current PowerShell executable for make wrapper launches."
|
||||
}
|
||||
|
||||
$wrapperPath = Join-Path $logDir "invoke-make.ps1"
|
||||
$escapedRepoRoot = $repoRoot -replace "'", "''"
|
||||
$wrapperTemplate = @'
|
||||
Set-StrictMode -Version Latest
|
||||
$ErrorActionPreference = "Stop"
|
||||
Set-Location -LiteralPath '{0}'
|
||||
function global:make {{
|
||||
{1}
|
||||
}}
|
||||
& make @args
|
||||
exit $LASTEXITCODE
|
||||
'@
|
||||
$wrapperScript = $wrapperTemplate -f $escapedRepoRoot, $command.Definition.TrimEnd()
|
||||
Set-Content -Path $wrapperPath -Value $wrapperScript -Encoding UTF8
|
||||
|
||||
return [PSCustomObject]@{
|
||||
FilePath = $shellPath
|
||||
ArgumentList = @("-NoLogo", "-NoProfile", "-File", $wrapperPath)
|
||||
}
|
||||
}
|
||||
|
||||
throw "code-change-verification: make must resolve to an application, script, alias, or function."
|
||||
}
|
||||
|
||||
$script:MakeInvocation = Resolve-MakeInvocation
|
||||
|
||||
function Invoke-MakeStep {
|
||||
param(
|
||||
[Parameter(Mandatory = $true)][string]$Step
|
||||
)
|
||||
|
||||
Write-Host "Running make $Step..."
|
||||
& $script:MakeInvocation.FilePath @($script:MakeInvocation.ArgumentList + $Step)
|
||||
|
||||
if ($LASTEXITCODE -ne 0) {
|
||||
Write-Host "code-change-verification: make $Step failed with exit code $LASTEXITCODE."
|
||||
return $LASTEXITCODE
|
||||
}
|
||||
|
||||
return 0
|
||||
}
|
||||
|
||||
function Start-MakeStep {
|
||||
param(
|
||||
[Parameter(Mandatory = $true)][string]$Step
|
||||
)
|
||||
|
||||
$stdoutLogPath = Join-Path $logDir "$Step.stdout.log"
|
||||
$stderrLogPath = Join-Path $logDir "$Step.stderr.log"
|
||||
Write-Host "Running make $Step..."
|
||||
$process = Start-Process -FilePath $script:MakeInvocation.FilePath -ArgumentList @($script:MakeInvocation.ArgumentList + $Step) -RedirectStandardOutput $stdoutLogPath -RedirectStandardError $stderrLogPath -PassThru
|
||||
$steps.Add([PSCustomObject]@{
|
||||
Name = $Step
|
||||
Process = $process
|
||||
StdoutLogPath = $stdoutLogPath
|
||||
StderrLogPath = $stderrLogPath
|
||||
StartTime = Get-Date
|
||||
})
|
||||
}
|
||||
|
||||
function Stop-RunningSteps {
|
||||
foreach ($step in $steps) {
|
||||
if ($null -eq $step.Process) {
|
||||
continue
|
||||
}
|
||||
|
||||
& taskkill /PID $step.Process.Id /T /F *> $null
|
||||
}
|
||||
|
||||
foreach ($step in $steps) {
|
||||
if ($null -eq $step.Process) {
|
||||
continue
|
||||
}
|
||||
|
||||
try {
|
||||
$step.Process.WaitForExit()
|
||||
} catch {
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
function Wait-ForParallelSteps {
|
||||
$pending = New-Object System.Collections.Generic.List[object]
|
||||
foreach ($step in $steps) {
|
||||
$pending.Add($step)
|
||||
}
|
||||
$nextHeartbeatAt = (Get-Date).AddSeconds($heartbeatIntervalSeconds)
|
||||
|
||||
while ($pending.Count -gt 0) {
|
||||
foreach ($step in @($pending)) {
|
||||
$step.Process.Refresh()
|
||||
if (-not $step.Process.HasExited) {
|
||||
continue
|
||||
}
|
||||
|
||||
$duration = [int]((Get-Date) - $step.StartTime).TotalSeconds
|
||||
if ($step.Process.ExitCode -eq 0) {
|
||||
Write-Host "make $($step.Name) passed in ${duration}s."
|
||||
[void]$pending.Remove($step)
|
||||
continue
|
||||
}
|
||||
|
||||
Write-Host "code-change-verification: make $($step.Name) failed with exit code $($step.Process.ExitCode) after ${duration}s."
|
||||
if (Test-Path $step.StderrLogPath) {
|
||||
Write-Host "--- $($step.Name) stderr log (last 80 lines) ---"
|
||||
Get-Content $step.StderrLogPath -Tail 80
|
||||
}
|
||||
if (Test-Path $step.StdoutLogPath) {
|
||||
Write-Host "--- $($step.Name) stdout log (last 80 lines) ---"
|
||||
Get-Content $step.StdoutLogPath -Tail 80
|
||||
}
|
||||
|
||||
Stop-RunningSteps
|
||||
return $step.Process.ExitCode
|
||||
}
|
||||
|
||||
if ($pending.Count -gt 0) {
|
||||
if ((Get-Date) -ge $nextHeartbeatAt) {
|
||||
$running = @()
|
||||
foreach ($step in $pending) {
|
||||
$elapsed = [int]((Get-Date) - $step.StartTime).TotalSeconds
|
||||
$running += "$($step.Name) (${elapsed}s)"
|
||||
}
|
||||
Write-Host ("code-change-verification: still running: " + ($running -join ", ") + ".")
|
||||
$nextHeartbeatAt = (Get-Date).AddSeconds($heartbeatIntervalSeconds)
|
||||
}
|
||||
Start-Sleep -Seconds 1
|
||||
}
|
||||
}
|
||||
|
||||
return 0
|
||||
}
|
||||
|
||||
$exitCode = 0
|
||||
|
||||
try {
|
||||
$exitCode = Invoke-MakeStep -Step "format"
|
||||
if ($exitCode -eq 0) {
|
||||
Write-Host "Running make lint, make typecheck, and make tests in parallel..."
|
||||
Start-MakeStep -Step "lint"
|
||||
Start-MakeStep -Step "typecheck"
|
||||
Start-MakeStep -Step "tests"
|
||||
|
||||
$exitCode = Wait-ForParallelSteps
|
||||
}
|
||||
} finally {
|
||||
Stop-RunningSteps
|
||||
Remove-Item $logDir -Recurse -Force -ErrorAction SilentlyContinue
|
||||
}
|
||||
|
||||
if ($exitCode -ne 0) {
|
||||
exit $exitCode
|
||||
}
|
||||
|
||||
Write-Host "code-change-verification: all commands passed."
|
||||
+390
@@ -0,0 +1,390 @@
|
||||
#!/usr/bin/env bash
|
||||
# Fail fast on any error or undefined variable.
|
||||
set -euo pipefail
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
if command -v git >/dev/null 2>&1; then
|
||||
REPO_ROOT="$(git -C "${SCRIPT_DIR}" rev-parse --show-toplevel 2>/dev/null || true)"
|
||||
fi
|
||||
REPO_ROOT="${REPO_ROOT:-$(cd "${SCRIPT_DIR}/../../../.." && pwd)}"
|
||||
|
||||
cd "${REPO_ROOT}"
|
||||
|
||||
LOG_DIR="$(mktemp -d "${TMPDIR:-/tmp}/code-change-verification.XXXXXX")"
|
||||
STATUS_PIPE="${LOG_DIR}/status.fifo"
|
||||
HEARTBEAT_INTERVAL_SECONDS="${CODE_CHANGE_VERIFICATION_HEARTBEAT_SECONDS:-10}"
|
||||
declare -a STEP_LAUNCHER=()
|
||||
declare -a STEP_PIDS=()
|
||||
declare -a STEP_NAMES=()
|
||||
declare -a STEP_LOGS=()
|
||||
declare -a STEP_STARTS=()
|
||||
RUNNING_STEPS=0
|
||||
EXIT_STATUS=0
|
||||
|
||||
resolve_executable_path() {
|
||||
local name="$1"
|
||||
type -P "${name}" 2>/dev/null || true
|
||||
}
|
||||
|
||||
configure_step_launcher() {
|
||||
local perl_path=""
|
||||
local python_path=""
|
||||
local uv_path=""
|
||||
|
||||
perl_path="$(resolve_executable_path perl)"
|
||||
if [ -n "${perl_path}" ]; then
|
||||
STEP_LAUNCHER=("${perl_path}" -MPOSIX=setsid -e 'setsid() or die $!; exec @ARGV')
|
||||
return 0
|
||||
fi
|
||||
|
||||
python_path="$(resolve_executable_path python3)"
|
||||
if [ -z "${python_path}" ]; then
|
||||
python_path="$(resolve_executable_path python)"
|
||||
fi
|
||||
if [ -n "${python_path}" ]; then
|
||||
STEP_LAUNCHER=("${python_path}" -c 'import os, sys; os.setsid(); os.execvp(sys.argv[1], sys.argv[1:])')
|
||||
return 0
|
||||
fi
|
||||
|
||||
uv_path="$(resolve_executable_path uv)"
|
||||
if [ -n "${uv_path}" ]; then
|
||||
STEP_LAUNCHER=("${uv_path}" run --no-sync python -c 'import os, sys; os.setsid(); os.execvp(sys.argv[1], sys.argv[1:])')
|
||||
return 0
|
||||
fi
|
||||
|
||||
echo "code-change-verification: perl, python3, python, or uv is required to manage parallel step process groups." >&2
|
||||
exit 1
|
||||
}
|
||||
|
||||
configure_step_launcher
|
||||
|
||||
mkfifo "${STATUS_PIPE}"
|
||||
exec 3<> "${STATUS_PIPE}"
|
||||
|
||||
cleanup() {
|
||||
local trap_status="$?"
|
||||
local status="${EXIT_STATUS}"
|
||||
|
||||
if [ "${status}" -eq 0 ]; then
|
||||
status="${trap_status}"
|
||||
fi
|
||||
|
||||
if [ "${#STEP_PIDS[@]}" -gt 0 ]; then
|
||||
stop_running_steps
|
||||
fi
|
||||
|
||||
exec 3>&- 3<&- || true
|
||||
rm -rf "${LOG_DIR}"
|
||||
exit "${status}"
|
||||
}
|
||||
|
||||
on_interrupt() {
|
||||
EXIT_STATUS=130
|
||||
exit 130
|
||||
}
|
||||
|
||||
on_terminate() {
|
||||
EXIT_STATUS=143
|
||||
exit 143
|
||||
}
|
||||
|
||||
stop_running_steps() {
|
||||
local pid=""
|
||||
|
||||
if [ "${#STEP_PIDS[@]}" -eq 0 ]; then
|
||||
return
|
||||
fi
|
||||
|
||||
for pid in "${STEP_PIDS[@]}"; do
|
||||
if [ -n "${pid}" ]; then
|
||||
kill -TERM -- "-${pid}" 2>/dev/null || true
|
||||
fi
|
||||
done
|
||||
|
||||
sleep 1
|
||||
|
||||
for pid in "${STEP_PIDS[@]}"; do
|
||||
if [ -n "${pid}" ]; then
|
||||
# A process group can remain alive after its leader exits, so escalate by group id unconditionally.
|
||||
kill -KILL -- "-${pid}" 2>/dev/null || true
|
||||
fi
|
||||
done
|
||||
|
||||
for pid in "${STEP_PIDS[@]}"; do
|
||||
if [ -n "${pid}" ]; then
|
||||
wait "${pid}" 2>/dev/null || true
|
||||
fi
|
||||
done
|
||||
|
||||
STEP_PIDS=()
|
||||
STEP_NAMES=()
|
||||
STEP_LOGS=()
|
||||
STEP_STARTS=()
|
||||
RUNNING_STEPS=0
|
||||
}
|
||||
|
||||
find_step_index() {
|
||||
local target_name="$1"
|
||||
local idx=""
|
||||
|
||||
for idx in "${!STEP_NAMES[@]}"; do
|
||||
if [ "${STEP_NAMES[$idx]}" = "${target_name}" ]; then
|
||||
echo "${idx}"
|
||||
return 0
|
||||
fi
|
||||
done
|
||||
|
||||
return 1
|
||||
}
|
||||
|
||||
clear_step() {
|
||||
local idx="$1"
|
||||
|
||||
STEP_PIDS[$idx]=""
|
||||
STEP_NAMES[$idx]=""
|
||||
STEP_LOGS[$idx]=""
|
||||
STEP_STARTS[$idx]=""
|
||||
RUNNING_STEPS=$((RUNNING_STEPS - 1))
|
||||
}
|
||||
|
||||
step_pid_is_alive() {
|
||||
local pid="$1"
|
||||
local state=""
|
||||
|
||||
if ! kill -0 "${pid}" 2>/dev/null; then
|
||||
return 1
|
||||
fi
|
||||
|
||||
state="$(ps -o stat= -p "${pid}" 2>/dev/null | tr -d '[:space:]')"
|
||||
case "${state}" in
|
||||
Z*|z*|"")
|
||||
return 1
|
||||
;;
|
||||
esac
|
||||
|
||||
return 0
|
||||
}
|
||||
|
||||
print_heartbeat() {
|
||||
local now
|
||||
local idx=""
|
||||
local name=""
|
||||
local start_time=""
|
||||
local elapsed=""
|
||||
local running=""
|
||||
|
||||
now=$(date +%s)
|
||||
|
||||
for idx in "${!STEP_NAMES[@]}"; do
|
||||
name="${STEP_NAMES[$idx]}"
|
||||
start_time="${STEP_STARTS[$idx]}"
|
||||
|
||||
if [ -z "${name}" ]; then
|
||||
continue
|
||||
fi
|
||||
|
||||
elapsed=$((now - start_time))
|
||||
if [ -n "${running}" ]; then
|
||||
running="${running}, "
|
||||
fi
|
||||
running="${running}${name} (${elapsed}s)"
|
||||
done
|
||||
|
||||
if [ -n "${running}" ]; then
|
||||
echo "code-change-verification: still running: ${running}."
|
||||
fi
|
||||
}
|
||||
|
||||
start_step() {
|
||||
local name="$1"
|
||||
shift
|
||||
local log_file="${LOG_DIR}/${name}.log"
|
||||
|
||||
echo "Running make ${name}..."
|
||||
: > "${log_file}"
|
||||
# Start each step in its own process group so fail-fast cleanup can stop pytest worker trees too.
|
||||
"${STEP_LAUNCHER[@]}" \
|
||||
bash -c '
|
||||
step_name="$1"
|
||||
log_file="$2"
|
||||
status_pipe="$3"
|
||||
shift 3
|
||||
|
||||
if "$@" >"$log_file" 2>&1; then
|
||||
status=0
|
||||
else
|
||||
status=$?
|
||||
fi
|
||||
|
||||
printf "%s\t%s\n" "$step_name" "$status" >"$status_pipe"
|
||||
exit "$status"
|
||||
' \
|
||||
bash "${name}" "${log_file}" "${STATUS_PIPE}" "$@" &
|
||||
|
||||
STEP_PIDS+=("$!")
|
||||
STEP_NAMES+=("${name}")
|
||||
STEP_LOGS+=("${log_file}")
|
||||
STEP_STARTS+=("$(date +%s)")
|
||||
RUNNING_STEPS=$((RUNNING_STEPS + 1))
|
||||
}
|
||||
|
||||
finish_step() {
|
||||
local name="$1"
|
||||
local status="$2"
|
||||
local idx=""
|
||||
local pid=""
|
||||
local log_file=""
|
||||
local start_time=""
|
||||
local now
|
||||
|
||||
idx="$(find_step_index "${name}")"
|
||||
pid="${STEP_PIDS[$idx]}"
|
||||
log_file="${STEP_LOGS[$idx]}"
|
||||
start_time="${STEP_STARTS[$idx]}"
|
||||
|
||||
now=$(date +%s)
|
||||
wait "${pid}" 2>/dev/null || true
|
||||
|
||||
if [ "${status}" -eq 0 ]; then
|
||||
clear_step "${idx}"
|
||||
echo "make ${name} passed in $((now - start_time))s."
|
||||
return 0
|
||||
fi
|
||||
|
||||
echo "code-change-verification: make ${name} failed with exit code ${status} after $((now - start_time))s." >&2
|
||||
echo "--- ${name} log (last 80 lines) ---" >&2
|
||||
tail -n 80 "${log_file}" >&2 || true
|
||||
stop_running_steps
|
||||
return "${status}"
|
||||
}
|
||||
|
||||
check_for_missing_reporters() {
|
||||
local idx=""
|
||||
local pid=""
|
||||
local name=""
|
||||
local log_file=""
|
||||
local start_time=""
|
||||
local now
|
||||
local step_status=0
|
||||
|
||||
for idx in "${!STEP_PIDS[@]}"; do
|
||||
pid="${STEP_PIDS[$idx]}"
|
||||
if [ -z "${pid}" ] || step_pid_is_alive "${pid}"; then
|
||||
continue
|
||||
fi
|
||||
|
||||
if try_finish_step_from_status_pipe 1; then
|
||||
if [ "${STATUS_PIPE_DRAINED}" -eq 1 ]; then
|
||||
return 0
|
||||
fi
|
||||
else
|
||||
step_status=$?
|
||||
return "${step_status}"
|
||||
fi
|
||||
|
||||
name="${STEP_NAMES[$idx]}"
|
||||
log_file="${STEP_LOGS[$idx]}"
|
||||
start_time="${STEP_STARTS[$idx]}"
|
||||
now=$(date +%s)
|
||||
wait "${pid}" 2>/dev/null || true
|
||||
|
||||
echo "code-change-verification: make ${name} exited before reporting completion status after $((now - start_time))s." >&2
|
||||
echo "--- ${name} log (last 80 lines) ---" >&2
|
||||
tail -n 80 "${log_file}" >&2 || true
|
||||
stop_running_steps
|
||||
return 1
|
||||
done
|
||||
|
||||
return 0
|
||||
}
|
||||
|
||||
STATUS_PIPE_DRAINED=0
|
||||
|
||||
try_finish_step_from_status_pipe() {
|
||||
local timeout="$1"
|
||||
local name=""
|
||||
local status=""
|
||||
local step_status=0
|
||||
|
||||
STATUS_PIPE_DRAINED=0
|
||||
if ! IFS=$'\t' read -r -t "${timeout}" name status <&3; then
|
||||
return 0
|
||||
fi
|
||||
|
||||
STATUS_PIPE_DRAINED=1
|
||||
finish_step "${name}" "${status}"
|
||||
step_status=$?
|
||||
if [ "${step_status}" -ne 0 ]; then
|
||||
return "${step_status}"
|
||||
fi
|
||||
|
||||
return 0
|
||||
}
|
||||
|
||||
wait_for_parallel_steps() {
|
||||
local name=""
|
||||
local status=""
|
||||
local step_status=""
|
||||
local next_heartbeat_at
|
||||
local now
|
||||
|
||||
next_heartbeat_at=$(( $(date +%s) + HEARTBEAT_INTERVAL_SECONDS ))
|
||||
|
||||
while [ "${RUNNING_STEPS}" -gt 0 ]; do
|
||||
if try_finish_step_from_status_pipe 1; then
|
||||
if [ "${STATUS_PIPE_DRAINED}" -eq 1 ]; then
|
||||
continue
|
||||
fi
|
||||
else
|
||||
step_status=$?
|
||||
if [ "${step_status}" -ne 0 ]; then
|
||||
return "${step_status}"
|
||||
fi
|
||||
continue
|
||||
fi
|
||||
|
||||
check_for_missing_reporters
|
||||
step_status=$?
|
||||
if [ "${step_status}" -ne 0 ]; then
|
||||
return "${step_status}"
|
||||
fi
|
||||
|
||||
now=$(date +%s)
|
||||
if [ "${now}" -ge "${next_heartbeat_at}" ]; then
|
||||
print_heartbeat
|
||||
next_heartbeat_at=$((now + HEARTBEAT_INTERVAL_SECONDS))
|
||||
fi
|
||||
done
|
||||
}
|
||||
|
||||
trap cleanup EXIT
|
||||
trap on_interrupt INT
|
||||
trap on_terminate TERM
|
||||
|
||||
echo "Running make format..."
|
||||
set +e
|
||||
make format
|
||||
EXIT_STATUS=$?
|
||||
set -e
|
||||
|
||||
if [ "${EXIT_STATUS}" -ne 0 ]; then
|
||||
exit "${EXIT_STATUS}"
|
||||
fi
|
||||
|
||||
echo "Running make lint, make typecheck, and make tests in parallel..."
|
||||
start_step "lint" make lint
|
||||
start_step "typecheck" make typecheck
|
||||
start_step "tests" make tests
|
||||
set +e
|
||||
wait_for_parallel_steps
|
||||
EXIT_STATUS=$?
|
||||
set -e
|
||||
|
||||
if [ "${EXIT_STATUS}" -ne 0 ]; then
|
||||
exit "${EXIT_STATUS}"
|
||||
fi
|
||||
|
||||
trap - EXIT INT TERM
|
||||
exec 3>&- 3<&-
|
||||
rm -rf "${LOG_DIR}"
|
||||
echo "code-change-verification: all commands passed."
|
||||
@@ -0,0 +1,76 @@
|
||||
---
|
||||
name: docs-sync
|
||||
description: Analyze main branch implementation and configuration to find missing, incorrect, or outdated documentation in docs/. Use when asked to audit doc coverage, sync docs with code, or propose doc updates/structure changes. Only update English docs under docs/** and never touch translated docs under docs/ja, docs/ko, or docs/zh. Provide a report and ask for approval before editing docs.
|
||||
---
|
||||
|
||||
# Docs Sync
|
||||
|
||||
## Overview
|
||||
|
||||
Identify doc coverage gaps and inaccuracies by comparing main branch features and configuration options against the current docs structure, then propose targeted improvements.
|
||||
|
||||
## Workflow
|
||||
|
||||
1. Confirm scope and base branch
|
||||
- Identify the current branch and default branch (usually `main`).
|
||||
- Prefer analyzing the current branch to keep work aligned with in-flight changes.
|
||||
- If the current branch is not `main`, analyze only the diff vs `main` to scope doc updates.
|
||||
- Avoid switching branches if it would disrupt local changes; use `git show main:<path>` or `git worktree add` when needed.
|
||||
|
||||
2. Build a feature inventory from the selected scope
|
||||
- If on `main`: inventory the full surface area and review docs comprehensively.
|
||||
- If not on `main`: inventory only changes vs `main` (feature additions/changes/removals).
|
||||
- Focus on user-facing behavior: public exports, configuration options, environment variables, CLI commands, default values, and documented runtime behaviors.
|
||||
- Capture evidence for each item (file path + symbol/setting).
|
||||
- Use targeted search to find option types and feature flags (for example: `rg "Settings"`, `rg "Config"`, `rg "os.environ"`, `rg "OPENAI_"`).
|
||||
- When the topic involves OpenAI platform features, invoke `$openai-knowledge` to pull current details from the OpenAI Developer Docs MCP server instead of guessing, while treating the SDK source code as the source of truth when discrepancies appear.
|
||||
|
||||
3. Doc-first pass: review existing pages
|
||||
- Walk each relevant page under `docs/` (excluding `docs/ja`, `docs/ko`, and `docs/zh`).
|
||||
- Identify missing mentions of important, supported options (opt-in flags, env vars), customization points, or new features from `src/agents/` and `examples/`.
|
||||
- Propose additions where users would reasonably expect to find them on that page.
|
||||
|
||||
4. Code-first pass: map features to docs
|
||||
- Review the current docs information architecture under `docs/` and `mkdocs.yml`.
|
||||
- Determine the best page/section for each feature based on existing patterns and the API reference structure under `docs/ref`.
|
||||
- Identify features that lack any doc page or have a page but no corresponding content.
|
||||
- Note when a structural adjustment would improve discoverability.
|
||||
- When improving `docs/ref/*` pages, treat the corresponding docstrings/comments in `src/` as the source of truth. Prefer updating those code comments so regenerated reference docs stay correct, instead of hand-editing the generated pages.
|
||||
|
||||
5. Detect gaps and inaccuracies
|
||||
- **Missing**: features/configs present in main but absent in docs.
|
||||
- **Incorrect/outdated**: names, defaults, or behaviors that diverge from main.
|
||||
- **Structural issues** (optional): pages overloaded, missing overviews, or mis-grouped topics.
|
||||
|
||||
6. Produce a Docs Sync Report and ask for approval
|
||||
- Provide a clear report with evidence, suggested doc locations, and proposed edits.
|
||||
- Ask the user whether to proceed with doc updates.
|
||||
|
||||
7. If approved, apply changes (English only)
|
||||
- Edit only English docs in `docs/**`.
|
||||
- Do **not** edit `docs/ja`, `docs/ko`, or `docs/zh`.
|
||||
- Keep changes aligned with the existing docs style and navigation.
|
||||
- Update `mkdocs.yml` when adding or renaming pages.
|
||||
- Build docs with `make build-docs` after edits to verify the docs site still builds.
|
||||
|
||||
## Output format
|
||||
|
||||
Use this template when reporting findings:
|
||||
|
||||
Docs Sync Report
|
||||
|
||||
- Doc-first findings
|
||||
- Page + missing content -> evidence + suggested insertion point
|
||||
- Code-first gaps
|
||||
- Feature + evidence -> suggested doc page/section (or missing page)
|
||||
- Incorrect or outdated docs
|
||||
- Doc file + issue + correct info + evidence
|
||||
- Structural suggestions (optional)
|
||||
- Proposed change + rationale
|
||||
- Proposed edits
|
||||
- Doc file -> concise change summary
|
||||
- Questions for the user
|
||||
|
||||
## References
|
||||
|
||||
- `references/doc-coverage-checklist.md`
|
||||
@@ -0,0 +1,4 @@
|
||||
interface:
|
||||
display_name: "Docs Sync"
|
||||
short_description: "Audit docs coverage and propose targeted updates"
|
||||
default_prompt: "Use $docs-sync to audit the current branch against docs/ and propose targeted documentation updates."
|
||||
@@ -0,0 +1,56 @@
|
||||
# Doc Coverage Checklist
|
||||
|
||||
Use this checklist to scan the selected scope (main = comprehensive, or current-branch diff) and validate documentation coverage.
|
||||
|
||||
## Feature inventory targets
|
||||
|
||||
- Public exports: classes, functions, types, and module entry points.
|
||||
- Configuration options: `*Settings` types, default config objects, and builder patterns.
|
||||
- Environment variables or runtime flags.
|
||||
- CLI commands, scripts, and example entry points that define supported usage.
|
||||
- User-facing behaviors: retry, timeouts, streaming, errors, logging, telemetry, and data handling.
|
||||
- Deprecations, removals, or renamed settings.
|
||||
|
||||
## Doc-first pass (page-by-page)
|
||||
|
||||
- Review each relevant English page (excluding `docs/ja`, `docs/ko`, and `docs/zh`).
|
||||
- Look for missing opt-in flags, env vars, or customization options that the page implies.
|
||||
- Add new features that belong on that page based on user intent and navigation.
|
||||
|
||||
## Code-first pass (feature inventory)
|
||||
|
||||
- Map features to the closest existing page based on the docs navigation in `mkdocs.yml`.
|
||||
- Prefer updating existing pages over creating new ones unless the topic is clearly new.
|
||||
- Use conceptual pages for cross-cutting concerns (auth, errors, streaming, tracing, tools).
|
||||
- Keep quick-start flows minimal; move advanced details into deeper pages.
|
||||
|
||||
## Evidence capture
|
||||
|
||||
- Record the main-branch file path and symbol/setting name.
|
||||
- Note defaults or behavior-critical details for accuracy checks.
|
||||
- Avoid large code dumps; a short identifier is enough.
|
||||
|
||||
## Red flags for outdated or incorrect docs
|
||||
|
||||
- Option names/types no longer exist or differ from code.
|
||||
- Default values or allowed ranges do not match implementation.
|
||||
- Features removed in code but still documented.
|
||||
- New behaviors introduced without corresponding docs updates.
|
||||
|
||||
## When to propose structural changes
|
||||
|
||||
- A page mixes unrelated audiences (quick-start + deep reference) without clear separation.
|
||||
- Multiple pages duplicate the same concept without cross-links.
|
||||
- New feature areas have no obvious home in the nav structure.
|
||||
|
||||
## Diff mode guidance (current branch vs main)
|
||||
|
||||
- Focus only on changed behavior: new exports/options, modified defaults, removed features, or renamed settings.
|
||||
- Use `git diff main...HEAD` (or equivalent) to constrain analysis.
|
||||
- Document removals explicitly so docs can be pruned if needed.
|
||||
|
||||
## Patch guidance
|
||||
|
||||
- Keep edits scoped and aligned with existing tone and format.
|
||||
- Update cross-links when moving or renaming sections.
|
||||
- Leave translated docs untouched; English-only updates.
|
||||
@@ -0,0 +1,90 @@
|
||||
---
|
||||
name: examples-auto-run
|
||||
description: Run python examples in auto mode with logging, rerun helpers, and background control.
|
||||
---
|
||||
|
||||
# examples-auto-run
|
||||
|
||||
## What it does
|
||||
|
||||
- Runs `uv run examples/run_examples.py` with:
|
||||
- Optional dependency extras enabled by default:
|
||||
`litellm`, `any-llm`, `sqlalchemy`, `redis`, `blaxel`, `modal`, `runloop`, and `temporal`.
|
||||
- `EXAMPLES_INTERACTIVE_MODE=auto` (auto-input/auto-approve).
|
||||
- Per-example logs under `.tmp/examples-start-logs/`.
|
||||
- Main summary log path passed via `--main-log` (also under `.tmp/examples-start-logs/`).
|
||||
- Generates a rerun list of failures at `.tmp/examples-rerun.txt` when `--write-rerun` is set.
|
||||
- Provides start/stop/status/logs/tail/collect/rerun helpers via `run.sh`.
|
||||
- Background option keeps the process running with a pidfile; `stop` cleans it up.
|
||||
|
||||
## Usage
|
||||
|
||||
```bash
|
||||
# Start (auto mode; interactive included by default)
|
||||
.agents/skills/examples-auto-run/scripts/run.sh start [extra args to run_examples.py]
|
||||
# Examples:
|
||||
.agents/skills/examples-auto-run/scripts/run.sh start --filter basic
|
||||
.agents/skills/examples-auto-run/scripts/run.sh start --include-server --include-audio
|
||||
|
||||
# Check status
|
||||
.agents/skills/examples-auto-run/scripts/run.sh status
|
||||
|
||||
# Stop running job
|
||||
.agents/skills/examples-auto-run/scripts/run.sh stop
|
||||
|
||||
# List logs
|
||||
.agents/skills/examples-auto-run/scripts/run.sh logs
|
||||
|
||||
# Tail latest log (or specify one)
|
||||
.agents/skills/examples-auto-run/scripts/run.sh tail
|
||||
.agents/skills/examples-auto-run/scripts/run.sh tail main_20260113-123000.log
|
||||
|
||||
# Collect rerun list from a main log (defaults to latest main_*.log)
|
||||
.agents/skills/examples-auto-run/scripts/run.sh collect
|
||||
|
||||
# Rerun only failed entries from rerun file (auto mode)
|
||||
.agents/skills/examples-auto-run/scripts/run.sh rerun
|
||||
```
|
||||
|
||||
## Codex execution environment
|
||||
|
||||
- When Codex invokes `start` or `rerun`, run the command outside the Codex sandbox by default
|
||||
(`sandbox_permissions=require_escalated`). Many examples start nested sandboxes, browsers, npm
|
||||
helpers, or cloud/local service processes; running from inside the Codex sandbox can produce
|
||||
environment-only failures such as `sandbox-exec: sandbox_apply: Operation not permitted`,
|
||||
Playwright cache permission errors, or npm cache permission errors.
|
||||
- Use sandboxed execution only when the user explicitly asks for it or when running a narrow dry-run
|
||||
/ log inspection command that does not execute examples.
|
||||
|
||||
## Defaults (overridable via env)
|
||||
|
||||
- `EXAMPLES_INTERACTIVE_MODE=auto`
|
||||
- `EXAMPLES_INCLUDE_INTERACTIVE=1`
|
||||
- `EXAMPLES_INCLUDE_SERVER=0`
|
||||
- `EXAMPLES_INCLUDE_AUDIO=0`
|
||||
- `EXAMPLES_INCLUDE_EXTERNAL=0`
|
||||
- `EXAMPLES_UV_EXTRAS="litellm any-llm sqlalchemy redis blaxel modal runloop temporal"` (set to an empty string to disable extras)
|
||||
- Auto-approvals in auto mode: `APPLY_PATCH_AUTO_APPROVE=1`, `SHELL_AUTO_APPROVE=1`, `AUTO_APPROVE_MCP=1`
|
||||
|
||||
## Log locations
|
||||
|
||||
- Main logs: `.tmp/examples-start-logs/main_*.log`
|
||||
- Per-example logs (from `run_examples.py`): `.tmp/examples-start-logs/<module_path>.log`
|
||||
- Rerun list: `.tmp/examples-rerun.txt`
|
||||
- Stdout logs: `.tmp/examples-start-logs/stdout_*.log`
|
||||
|
||||
## Notes
|
||||
|
||||
- The runner delegates to `uv run --extra ... examples/run_examples.py`, which already writes per-example logs and supports `--collect`, `--rerun-file`, and `--print-auto-skip`.
|
||||
- `start` uses `--write-rerun` so failures are captured automatically.
|
||||
- If `.tmp/examples-rerun.txt` exists and is non-empty, invoking the skill with no args runs `rerun` by default.
|
||||
|
||||
## Behavioral validation (Codex/LLM responsibility)
|
||||
|
||||
The runner does not perform any automated behavioral validation. After every foreground `start` or `rerun`, **Codex must manually validate** all exit-0 entries:
|
||||
|
||||
1. Read the example source (and comments) to infer intended flow, tools used, and expected key outputs.
|
||||
2. Open the matching per-example log under `.tmp/examples-start-logs/`.
|
||||
3. Confirm the intended actions/results occurred; flag omissions or divergences.
|
||||
4. Do this for **all passed examples**, not just a sample.
|
||||
5. Report immediately after the run with concise citations to the exact log lines that justify the validation.
|
||||
@@ -0,0 +1,4 @@
|
||||
interface:
|
||||
display_name: "Examples Auto Run"
|
||||
short_description: "Run examples in auto mode with logs and rerun helpers"
|
||||
default_prompt: "Use $examples-auto-run to run the repo examples in auto mode, collect logs, and summarize any failures."
|
||||
+236
@@ -0,0 +1,236 @@
|
||||
#!/usr/bin/env bash
|
||||
set -euo pipefail
|
||||
|
||||
ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/../../../.." && pwd)"
|
||||
PID_FILE="$ROOT/.tmp/examples-auto-run.pid"
|
||||
LOG_DIR="$ROOT/.tmp/examples-start-logs"
|
||||
RERUN_FILE="$ROOT/.tmp/examples-rerun.txt"
|
||||
DEFAULT_UV_EXTRAS="litellm any-llm sqlalchemy redis blaxel modal runloop temporal"
|
||||
|
||||
build_uv_prefix() {
|
||||
UV_RUN=(uv run)
|
||||
local extras_value
|
||||
if [[ -n "${EXAMPLES_UV_EXTRAS+x}" ]]; then
|
||||
extras_value="$EXAMPLES_UV_EXTRAS"
|
||||
else
|
||||
extras_value="$DEFAULT_UV_EXTRAS"
|
||||
fi
|
||||
|
||||
local extra
|
||||
for extra in $extras_value; do
|
||||
UV_RUN+=(--extra "$extra")
|
||||
done
|
||||
export EXAMPLES_UV_EXTRAS="$extras_value"
|
||||
}
|
||||
|
||||
ensure_dirs() {
|
||||
mkdir -p "$LOG_DIR" "$ROOT/.tmp"
|
||||
}
|
||||
|
||||
is_running() {
|
||||
local pid="$1"
|
||||
[[ -n "$pid" ]] && ps -p "$pid" >/dev/null 2>&1
|
||||
}
|
||||
|
||||
cmd_start() {
|
||||
ensure_dirs
|
||||
local background=0
|
||||
if [[ "${1:-}" == "--background" ]]; then
|
||||
background=1
|
||||
shift
|
||||
fi
|
||||
|
||||
local ts main_log stdout_log
|
||||
ts="$(date +%Y%m%d-%H%M%S)"
|
||||
main_log="$LOG_DIR/main_${ts}.log"
|
||||
stdout_log="$LOG_DIR/stdout_${ts}.log"
|
||||
|
||||
build_uv_prefix
|
||||
local run_cmd=(
|
||||
"${UV_RUN[@]}" examples/run_examples.py
|
||||
--auto-mode
|
||||
--write-rerun
|
||||
--main-log "$main_log"
|
||||
--logs-dir "$LOG_DIR"
|
||||
)
|
||||
|
||||
if [[ "$background" -eq 1 ]]; then
|
||||
if [[ -f "$PID_FILE" ]]; then
|
||||
local pid
|
||||
pid="$(cat "$PID_FILE" 2>/dev/null || true)"
|
||||
if is_running "$pid"; then
|
||||
echo "examples/run_examples.py already running (pid=$pid)."
|
||||
exit 1
|
||||
fi
|
||||
fi
|
||||
(
|
||||
trap '' HUP
|
||||
export EXAMPLES_INTERACTIVE_MODE="${EXAMPLES_INTERACTIVE_MODE:-auto}"
|
||||
export APPLY_PATCH_AUTO_APPROVE="${APPLY_PATCH_AUTO_APPROVE:-1}"
|
||||
export SHELL_AUTO_APPROVE="${SHELL_AUTO_APPROVE:-1}"
|
||||
export AUTO_APPROVE_MCP="${AUTO_APPROVE_MCP:-1}"
|
||||
export EXAMPLES_INCLUDE_INTERACTIVE="${EXAMPLES_INCLUDE_INTERACTIVE:-1}"
|
||||
export EXAMPLES_INCLUDE_SERVER="${EXAMPLES_INCLUDE_SERVER:-0}"
|
||||
export EXAMPLES_INCLUDE_AUDIO="${EXAMPLES_INCLUDE_AUDIO:-0}"
|
||||
export EXAMPLES_INCLUDE_EXTERNAL="${EXAMPLES_INCLUDE_EXTERNAL:-0}"
|
||||
cd "$ROOT"
|
||||
exec "${run_cmd[@]}" "$@" > >(tee "$stdout_log") 2>&1
|
||||
) &
|
||||
local pid=$!
|
||||
echo "$pid" >"$PID_FILE"
|
||||
echo "Started run_examples.py (pid=$pid)"
|
||||
echo "Main log: $main_log"
|
||||
echo "Stdout log: $stdout_log"
|
||||
echo "Run '.agents/skills/examples-auto-run/scripts/run.sh validate \"$main_log\"' after it finishes."
|
||||
return 0
|
||||
fi
|
||||
|
||||
export EXAMPLES_INTERACTIVE_MODE="${EXAMPLES_INTERACTIVE_MODE:-auto}"
|
||||
export APPLY_PATCH_AUTO_APPROVE="${APPLY_PATCH_AUTO_APPROVE:-1}"
|
||||
export SHELL_AUTO_APPROVE="${SHELL_AUTO_APPROVE:-1}"
|
||||
export AUTO_APPROVE_MCP="${AUTO_APPROVE_MCP:-1}"
|
||||
export EXAMPLES_INCLUDE_INTERACTIVE="${EXAMPLES_INCLUDE_INTERACTIVE:-1}"
|
||||
export EXAMPLES_INCLUDE_SERVER="${EXAMPLES_INCLUDE_SERVER:-0}"
|
||||
export EXAMPLES_INCLUDE_AUDIO="${EXAMPLES_INCLUDE_AUDIO:-0}"
|
||||
export EXAMPLES_INCLUDE_EXTERNAL="${EXAMPLES_INCLUDE_EXTERNAL:-0}"
|
||||
cd "$ROOT"
|
||||
set +e
|
||||
"${run_cmd[@]}" "$@" 2>&1 | tee "$stdout_log"
|
||||
local run_status=${PIPESTATUS[0]}
|
||||
set -e
|
||||
return "$run_status"
|
||||
}
|
||||
|
||||
cmd_stop() {
|
||||
if [[ ! -f "$PID_FILE" ]]; then
|
||||
echo "No pid file; nothing to stop."
|
||||
return 0
|
||||
fi
|
||||
local pid
|
||||
pid="$(cat "$PID_FILE" 2>/dev/null || true)"
|
||||
if [[ -z "$pid" ]]; then
|
||||
rm -f "$PID_FILE"
|
||||
echo "Pid file empty; cleaned."
|
||||
return 0
|
||||
fi
|
||||
if ! is_running "$pid"; then
|
||||
rm -f "$PID_FILE"
|
||||
echo "Process $pid not running; cleaned pid file."
|
||||
return 0
|
||||
fi
|
||||
echo "Stopping pid $pid ..."
|
||||
kill "$pid" 2>/dev/null || true
|
||||
sleep 1
|
||||
if is_running "$pid"; then
|
||||
echo "Sending SIGKILL to $pid ..."
|
||||
kill -9 "$pid" 2>/dev/null || true
|
||||
fi
|
||||
rm -f "$PID_FILE"
|
||||
echo "Stopped."
|
||||
}
|
||||
|
||||
cmd_status() {
|
||||
if [[ -f "$PID_FILE" ]]; then
|
||||
local pid
|
||||
pid="$(cat "$PID_FILE" 2>/dev/null || true)"
|
||||
if is_running "$pid"; then
|
||||
echo "Running (pid=$pid)"
|
||||
return 0
|
||||
fi
|
||||
fi
|
||||
echo "Not running."
|
||||
}
|
||||
|
||||
cmd_logs() {
|
||||
ensure_dirs
|
||||
ls -1t "$LOG_DIR"
|
||||
}
|
||||
|
||||
cmd_tail() {
|
||||
ensure_dirs
|
||||
local file="${1:-}"
|
||||
if [[ -z "$file" ]]; then
|
||||
file="$(ls -1t "$LOG_DIR" | head -n1)"
|
||||
fi
|
||||
if [[ -z "$file" ]]; then
|
||||
echo "No log files yet."
|
||||
exit 1
|
||||
fi
|
||||
tail -f "$LOG_DIR/$file"
|
||||
}
|
||||
|
||||
collect_rerun() {
|
||||
ensure_dirs
|
||||
local log_file="${1:-}"
|
||||
if [[ -z "$log_file" ]]; then
|
||||
log_file="$(ls -1t "$LOG_DIR"/main_*.log 2>/dev/null | head -n1)"
|
||||
fi
|
||||
if [[ -z "$log_file" ]] || [[ ! -f "$log_file" ]]; then
|
||||
echo "No main log file found."
|
||||
exit 1
|
||||
fi
|
||||
cd "$ROOT"
|
||||
build_uv_prefix
|
||||
"${UV_RUN[@]}" examples/run_examples.py --collect "$log_file" --output "$RERUN_FILE"
|
||||
}
|
||||
|
||||
cmd_rerun() {
|
||||
ensure_dirs
|
||||
local file="${1:-$RERUN_FILE}"
|
||||
if [[ ! -s "$file" ]]; then
|
||||
echo "Rerun list is empty: $file"
|
||||
exit 0
|
||||
fi
|
||||
local ts main_log stdout_log
|
||||
ts="$(date +%Y%m%d-%H%M%S)"
|
||||
main_log="$LOG_DIR/main_${ts}.log"
|
||||
stdout_log="$LOG_DIR/stdout_${ts}.log"
|
||||
cd "$ROOT"
|
||||
export EXAMPLES_INTERACTIVE_MODE="${EXAMPLES_INTERACTIVE_MODE:-auto}"
|
||||
export APPLY_PATCH_AUTO_APPROVE="${APPLY_PATCH_AUTO_APPROVE:-1}"
|
||||
export SHELL_AUTO_APPROVE="${SHELL_AUTO_APPROVE:-1}"
|
||||
export AUTO_APPROVE_MCP="${AUTO_APPROVE_MCP:-1}"
|
||||
build_uv_prefix
|
||||
set +e
|
||||
"${UV_RUN[@]}" examples/run_examples.py --auto-mode --rerun-file "$file" --write-rerun --main-log "$main_log" --logs-dir "$LOG_DIR" 2>&1 | tee "$stdout_log"
|
||||
local run_status=${PIPESTATUS[0]}
|
||||
set -e
|
||||
return "$run_status"
|
||||
}
|
||||
|
||||
usage() {
|
||||
cat <<'EOF'
|
||||
Usage: run.sh <start|stop|status|logs|tail|collect|rerun> [args...]
|
||||
|
||||
Commands:
|
||||
start [--filter ... | other args] Run examples in auto mode (foreground). Pass --background to run detached.
|
||||
stop Kill the running auto-run (if any).
|
||||
status Show whether it is running.
|
||||
logs List log files (.tmp/examples-start-logs).
|
||||
tail [logfile] Tail the latest (or specified) log.
|
||||
collect [main_log] Parse a main log and write failed examples to .tmp/examples-rerun.txt.
|
||||
rerun [rerun_file] Run only the examples listed in .tmp/examples-rerun.txt.
|
||||
|
||||
Environment overrides:
|
||||
EXAMPLES_INTERACTIVE_MODE (default auto)
|
||||
EXAMPLES_INCLUDE_SERVER/INTERACTIVE/AUDIO/EXTERNAL (defaults: 0/1/0/0)
|
||||
EXAMPLES_UV_EXTRAS (default: litellm any-llm sqlalchemy redis blaxel modal runloop; set empty to disable)
|
||||
APPLY_PATCH_AUTO_APPROVE, SHELL_AUTO_APPROVE, AUTO_APPROVE_MCP (default 1 in auto mode)
|
||||
EOF
|
||||
}
|
||||
|
||||
default_cmd="start"
|
||||
if [[ $# -eq 0 && -s "$RERUN_FILE" ]]; then
|
||||
default_cmd="rerun"
|
||||
fi
|
||||
|
||||
case "${1:-$default_cmd}" in
|
||||
start) shift || true; cmd_start "$@" ;;
|
||||
stop) shift || true; cmd_stop ;;
|
||||
status) shift || true; cmd_status ;;
|
||||
logs) shift || true; cmd_logs ;;
|
||||
tail) shift; cmd_tail "${1:-}" ;;
|
||||
collect) shift || true; collect_rerun "${1:-}" ;;
|
||||
rerun) shift || true; cmd_rerun "${1:-}" ;;
|
||||
*) usage; exit 1 ;;
|
||||
esac
|
||||
@@ -0,0 +1,126 @@
|
||||
---
|
||||
name: final-release-review
|
||||
description: Perform a release-readiness review by locating the previous release tag from remote tags and auditing the diff (e.g., v1.2.3...<commit>) for breaking changes, regressions, improvement opportunities, and risks before releasing openai-agents-python.
|
||||
---
|
||||
|
||||
# Final Release Review
|
||||
|
||||
## Purpose
|
||||
|
||||
Use this skill when validating the latest release candidate commit (default tip of `origin/main`) for release. It guides you to fetch remote tags, pick the previous release tag, and thoroughly inspect the `BASE_TAG...TARGET` diff for breaking changes, introduced bugs/regressions, improvement opportunities, and release risks.
|
||||
|
||||
The review must be stable and actionable: avoid variance between runs by using explicit gate rules, and never produce a `BLOCKED` call without concrete evidence and clear unblock actions.
|
||||
|
||||
## Quick start
|
||||
|
||||
1. Ensure repository root: `pwd` → `path-to-workspace/openai-agents-python`.
|
||||
2. Sync tags and pick base (default `v*`):
|
||||
```bash
|
||||
BASE_TAG="$(.agents/skills/final-release-review/scripts/find_latest_release_tag.sh origin 'v*')"
|
||||
```
|
||||
3. Choose target commit (default tip of `origin/main`, ensure fresh): `git fetch origin main --prune` then `TARGET="$(git rev-parse origin/main)"`.
|
||||
4. Snapshot scope:
|
||||
```bash
|
||||
git diff --stat "${BASE_TAG}"..."${TARGET}"
|
||||
git diff --dirstat=files,0 "${BASE_TAG}"..."${TARGET}"
|
||||
git log --oneline --reverse "${BASE_TAG}".."${TARGET}"
|
||||
git diff --name-status "${BASE_TAG}"..."${TARGET}"
|
||||
```
|
||||
5. Deep review using `references/review-checklist.md` to spot breaking changes, regressions, and improvement chances.
|
||||
6. Capture findings and call the release gate: ship/block with conditions; propose focused tests for risky areas.
|
||||
|
||||
## Deterministic gate policy
|
||||
|
||||
- Default to **🟢 GREEN LIGHT TO SHIP** unless at least one blocking trigger below is satisfied.
|
||||
- Use **🔴 BLOCKED** only when you can cite concrete release-blocking evidence and provide actionable unblock steps.
|
||||
- Blocking triggers (at least one required for `BLOCKED`):
|
||||
- A confirmed regression or bug introduced in `BASE...TARGET` (for example, failing targeted test, incompatible behavior in diff, or removed behavior without fallback).
|
||||
- A confirmed breaking public API/protocol/config change with missing or mismatched versioning and no migration path (for example, patch release for a breaking change).
|
||||
- A concrete data-loss, corruption, or security-impacting change with unresolved mitigation.
|
||||
- A release-critical packaging/build/runtime path is broken by the diff (not speculative).
|
||||
- Non-blocking by itself:
|
||||
- Large diff size, broad refactor, or many touched files.
|
||||
- "Could regress" risk statements without concrete evidence.
|
||||
- Not running tests locally.
|
||||
- If evidence is incomplete, issue **🟢 GREEN LIGHT TO SHIP** with targeted validation follow-ups instead of `BLOCKED`.
|
||||
|
||||
## Workflow
|
||||
|
||||
- **Prepare**
|
||||
- Run the quick-start tag command to ensure you use the latest remote tag. If the tag pattern differs, override the pattern argument (e.g., `'*.*.*'`).
|
||||
- If the user specifies a base tag, prefer it but still fetch remote tags first.
|
||||
- Keep the working tree clean to avoid diff noise.
|
||||
- **Assumptions**
|
||||
- Assume the target commit (default `origin/main` tip) has already passed `$code-change-verification` in CI unless the user says otherwise.
|
||||
- Do not block a release solely because you did not run tests locally; focus on concrete behavioral or API risks.
|
||||
- Release policy: routine releases use patch versions; use minor only for breaking changes or major feature additions. Major versions are reserved until the 1.0 release.
|
||||
- **Map the diff**
|
||||
- Use `--stat`, `--dirstat`, and `--name-status` outputs to spot hot directories and file types.
|
||||
- For suspicious files, prefer `git diff --word-diff BASE...TARGET -- <path>`.
|
||||
- Note any deleted or newly added tests, config, migrations, or scripts.
|
||||
- **Analyze risk**
|
||||
- Walk through the categories in `references/review-checklist.md` (breaking changes, regression clues, improvement opportunities).
|
||||
- When you suspect a risk, cite the specific file/commit and explain the behavioral impact.
|
||||
- For every finding, include all of: `Evidence`, `Impact`, and `Action`.
|
||||
- Severity calibration:
|
||||
- **🟢 LOW**: low blast radius or clearly covered behavior; no release gate impact.
|
||||
- **🟡 MODERATE**: plausible user-facing regression signal; needs validation but not a confirmed blocker.
|
||||
- **🔴 HIGH**: confirmed or strongly evidenced release-blocking issue.
|
||||
- Suggest minimal, high-signal validation commands (targeted tests or linters) instead of generic reruns when time is tight.
|
||||
- Breaking changes do not automatically require a BLOCKED release call when they are already covered by an appropriate version bump and migration/upgrade notes; only block when the bump is missing/mismatched (e.g., patch bump) or when the breaking change introduces unresolved risk.
|
||||
- **Form a recommendation**
|
||||
- State BASE_TAG and TARGET explicitly.
|
||||
- Provide a concise diff summary (key directories/files and counts).
|
||||
- List: breaking-change candidates, probable regressions/bugs, improvement opportunities, missing release notes/migrations.
|
||||
- Recommend ship/block and the exact checks needed to unblock if blocking. If a breaking change is properly versioned (minor/major), you may still recommend a GREEN LIGHT TO SHIP while calling out the change. Use emoji and boldface in the release call to make the gate obvious.
|
||||
- If you cannot provide a concrete unblock checklist item, do not use `BLOCKED`.
|
||||
|
||||
## Output format (required)
|
||||
|
||||
All output must be in English.
|
||||
|
||||
Use the following report structure in every response produced by this skill. Be proactive and decisive: make a clear ship/block call near the top, and assign an explicit risk level (LOW/MODERATE/HIGH) to each finding with a short impact statement. Avoid overly cautious hedging when the risk is low and tests passed.
|
||||
|
||||
Always use the fixed repository URL in the Diff section (`https://github.com/openai/openai-agents-python/compare/...`). Do not use `${GITHUB_REPOSITORY}` or any other template variable. Format risk levels as bold emoji labels: **🟢 LOW**, **🟡 MODERATE**, **🔴 HIGH**.
|
||||
|
||||
Every risk finding must contain an actionable next step. If the report uses `**🔴 BLOCKED**`, include an `Unblock checklist` section with at least one concrete command/task and a pass condition.
|
||||
|
||||
```
|
||||
### Release readiness review (<tag> -> TARGET <ref>)
|
||||
|
||||
This is a release readiness report done by `$final-release-review` skill.
|
||||
|
||||
### Diff
|
||||
|
||||
https://github.com/openai/openai-agents-python/compare/<tag>...<target-commit>
|
||||
|
||||
### Release call:
|
||||
**<🟢 GREEN LIGHT TO SHIP | 🔴 BLOCKED>** <one-line rationale>
|
||||
|
||||
### Scope summary:
|
||||
- <N files changed (+A/-D); key areas touched: ...>
|
||||
|
||||
### Risk assessment (ordered by impact):
|
||||
1) **<Finding title>**
|
||||
- Risk: **<🟢 LOW | 🟡 MODERATE | 🔴 HIGH>**. <Impact statement in one sentence.>
|
||||
- Evidence: <specific diff/test/commit signal; avoid generic statements>
|
||||
- Files: <path(s)>
|
||||
- Action: <concrete next step command/task with pass criteria>
|
||||
2) ...
|
||||
|
||||
### Unblock checklist (required when Release call is BLOCKED):
|
||||
1. [ ] <concrete check/fix>
|
||||
- Exit criteria: <what must be true to unblock>
|
||||
2. ...
|
||||
|
||||
### Notes:
|
||||
- <working tree status, tag/target assumptions, or re-run guidance>
|
||||
```
|
||||
|
||||
If no risks are found, include a “No material risks identified” line under Risk assessment and still provide a ship call. If you did not run local verification, do not add a verification status section or use it as a release blocker; note any assumptions briefly in Notes.
|
||||
If the report is not blocked, omit the `Unblock checklist` section.
|
||||
|
||||
### Resources
|
||||
|
||||
- `scripts/find_latest_release_tag.sh`: Fetches remote tags and returns the newest tag matching a pattern (default `v*`).
|
||||
- `references/review-checklist.md`: Detailed signals and commands for spotting breaking changes, regressions, and release polish gaps.
|
||||
@@ -0,0 +1,4 @@
|
||||
interface:
|
||||
display_name: "Final Release Review"
|
||||
short_description: "Audit a release candidate against the previous tag"
|
||||
default_prompt: "Use $final-release-review to audit the release candidate diff against the previous release tag and call the ship/block gate."
|
||||
@@ -0,0 +1,65 @@
|
||||
# Release Diff Review Checklist
|
||||
|
||||
## Quick commands
|
||||
|
||||
- Sync tags: `git fetch origin --tags --prune`.
|
||||
- Identify latest release tag (default pattern `v*`): `git tag -l 'v*' --sort=-v:refname | head -n1` or use `.agents/skills/final-release-review/scripts/find_latest_release_tag.sh`.
|
||||
- Generate overview: `git diff --stat BASE...TARGET`, `git diff --dirstat=files,0 BASE...TARGET`, `git log --oneline --reverse BASE..TARGET`.
|
||||
- Inspect risky files quickly: `git diff --name-status BASE...TARGET`, `git diff --word-diff BASE...TARGET -- <path>`.
|
||||
|
||||
## Gate decision matrix
|
||||
|
||||
- Choose `🟢 GREEN LIGHT TO SHIP` when no concrete blocking trigger is found.
|
||||
- Choose `🔴 BLOCKED` only when at least one blocking trigger has concrete evidence and a defined unblock action.
|
||||
- Blocking triggers:
|
||||
- Confirmed regression/bug introduced in the diff.
|
||||
- Confirmed breaking public API/protocol/config change with missing or mismatched versioning/migration path.
|
||||
- Concrete data-loss/corruption/security-impacting issue with unresolved mitigation.
|
||||
- Release-critical build/package/runtime break introduced by the diff.
|
||||
- Non-blocking by itself:
|
||||
- Large refactor or high file count.
|
||||
- Speculative risk without evidence.
|
||||
- Not running tests locally.
|
||||
- If uncertain, keep gate green and provide focused follow-up checks.
|
||||
|
||||
## Actionability contract
|
||||
|
||||
- Every risk finding should include:
|
||||
- `Evidence`: specific file/commit/diff/test signal.
|
||||
- `Impact`: one-sentence user or runtime effect.
|
||||
- `Action`: concrete command/task with pass criteria.
|
||||
- A `BLOCKED` report must contain an `Unblock checklist` with at least one executable item.
|
||||
- If no executable unblock item exists, do not block; downgrade to green with follow-up checks.
|
||||
|
||||
## Breaking change signals
|
||||
|
||||
- Public API surface: removed/renamed modules, classes, functions, or re-exports; changed parameters/return types, default values changed, new required options, stricter validation.
|
||||
- Protocol/schema: request/response fields added/removed/renamed, enum changes, JSON shape changes, ID formats, pagination defaults.
|
||||
- Config/CLI/env: renamed flags, default behavior flips, removed fallbacks, environment variable changes, logging levels tightened.
|
||||
- Dependencies/platform: Python version requirement changes, dependency major bumps, `pyproject.toml`/`uv.lock` changes, removed or renamed extras.
|
||||
- Persistence/data: migration scripts missing, data model changes, stored file formats, cache keys altered without invalidation.
|
||||
- Docs/examples drift: examples still reflect old behavior or lack migration note.
|
||||
|
||||
## Regression risk clues
|
||||
|
||||
- Large refactors with light test deltas or deleted tests; new `skip`/`todo` markers.
|
||||
- Concurrency/timing: new async flows, asyncio event-loop changes, retries, timeouts, debounce/caching changes, race-prone patterns.
|
||||
- Error handling: catch blocks removed, swallowed errors, broader catch-all added without logging, stricter throws without caller updates.
|
||||
- Stateful components: mutable shared state, global singletons, lifecycle changes (init/teardown), resource cleanup removal.
|
||||
- Third-party changes: swapped core libraries, feature flags toggled, observability removed or gated.
|
||||
|
||||
## Improvement opportunities
|
||||
|
||||
- Missing coverage for new code paths; add focused tests.
|
||||
- Performance: obvious N+1 loops, repeated I/O without caching, excessive serialization.
|
||||
- Developer ergonomics: unclear naming, missing inline docs for public APIs, missing examples for new features.
|
||||
- Release hygiene: add migration/upgrade note when behavior changes; ensure changelog/notes capture user-facing shifts.
|
||||
|
||||
## Evidence to capture in the review output
|
||||
|
||||
- BASE tag and TARGET ref used for the diff; confirm tags fetched.
|
||||
- High-level diff stats and key directories touched.
|
||||
- Concrete files/commits that indicate breaking changes or risk, with brief rationale.
|
||||
- Tests or commands suggested to validate suspected risks (include pass criteria).
|
||||
- Explicit release gate call (ship/block) with conditions to unblock.
|
||||
- `Unblock checklist` section when (and only when) gate is `BLOCKED`.
|
||||
@@ -0,0 +1,17 @@
|
||||
#!/usr/bin/env bash
|
||||
set -euo pipefail
|
||||
|
||||
remote="${1:-origin}"
|
||||
pattern="${2:-v*}"
|
||||
|
||||
# Sync tags from the remote to ensure the latest release tag is available locally.
|
||||
git fetch "$remote" --tags --prune --quiet
|
||||
|
||||
latest_tag=$(git tag -l "$pattern" --sort=-v:refname | head -n1)
|
||||
|
||||
if [[ -z "$latest_tag" ]]; then
|
||||
echo "No tags found matching pattern '$pattern' after fetching from $remote." >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
echo "$latest_tag"
|
||||
@@ -0,0 +1,61 @@
|
||||
---
|
||||
name: implementation-strategy
|
||||
description: Decide how to implement runtime and API changes in openai-agents-python before editing code. Use when a task changes exported APIs, runtime behavior, serialized state, tests, or docs and you need to choose the compatibility boundary, whether shims or migrations are warranted, and when unreleased interfaces can be rewritten directly.
|
||||
---
|
||||
|
||||
# Implementation Strategy
|
||||
|
||||
## Overview
|
||||
|
||||
Use this skill before editing code when the task changes runtime behavior or anything that might look like a compatibility concern. The goal is to keep implementations simple while protecting real released contracts.
|
||||
|
||||
## Quick start
|
||||
|
||||
1. Identify the surface you are changing: released public API, unreleased branch-local API, internal helper, persisted schema, wire protocol, CLI/config/env surface, or docs/examples only.
|
||||
2. Determine the latest release boundary from `origin` first, and only fall back to local tags when remote tags are unavailable:
|
||||
```bash
|
||||
BASE_TAG="$(.agents/skills/final-release-review/scripts/find_latest_release_tag.sh origin 'v*' 2>/dev/null || git tag -l 'v*' --sort=-v:refname | head -n1)"
|
||||
echo "$BASE_TAG"
|
||||
```
|
||||
3. Judge breaking-change risk against that latest release tag, not against unreleased branch churn or post-tag changes already on `main`. If the command fell back to local tags, treat the result as potentially stale and say so.
|
||||
4. Prefer the simplest implementation that satisfies the current task. Update callers, tests, docs, and examples directly instead of preserving superseded unreleased interfaces.
|
||||
5. Add a compatibility layer only when there is a concrete released consumer, an otherwise supported durable external state boundary that requires it, or when the user explicitly asks for a migration path.
|
||||
|
||||
## Compatibility boundary rules
|
||||
|
||||
- Released public API or documented external behavior: preserve compatibility or provide an explicit migration path.
|
||||
- Persisted schema, serialized state, wire protocol, CLI flags, environment variables, and externally consumed config: treat as compatibility-sensitive when they are part of the latest release or when the repo explicitly intends to preserve them across commits, processes, or machines.
|
||||
- Python-specific durable surfaces such as `RunState`, session persistence, exported dataclass constructor order, and documented model/provider configuration should be treated as compatibility-sensitive when they were part of the latest release tag or are explicitly supported as a shared durability boundary.
|
||||
- Interface changes introduced only on the current branch: not a compatibility target. Rewrite them directly.
|
||||
- Interface changes present on `main` but added after the latest release tag: not a semver breaking change by themselves. Rewrite them directly unless they already define a released or explicitly supported durable external state boundary.
|
||||
- Internal helpers, private types, same-branch tests, fixtures, and examples: update them directly instead of adding adapters.
|
||||
- Unreleased persisted schema versions on `main` may be renumbered or squashed before release when intermediate snapshots are intentionally unsupported. When you do that, update the support set and tests together so the boundary is explicit.
|
||||
|
||||
## Default implementation stance
|
||||
|
||||
- Prefer deletion or replacement over aliases, overloads, shims, feature flags, and dual-write logic when the old shape is unreleased.
|
||||
- Do not preserve a confusing abstraction just because it exists in the current branch diff.
|
||||
- If review feedback claims a change is breaking, verify it against the latest release tag and actual external impact before accepting the feedback.
|
||||
- If a change truly crosses the latest released contract boundary, call that out explicitly in the ExecPlan, release notes context, and user-facing summary.
|
||||
|
||||
## SDK-specific decision rules
|
||||
|
||||
- When unsupported OpenAI API or provider-adapter behavior already has a released default path, avoid turning it into a default hard error unless the latest release boundary justifies that break. Prefer an opt-in strict mode such as `strict_feature_validation=True`, while keeping the default path compatible through warning, ignoring unsupported data, or a clearly non-empty placeholder.
|
||||
- For OpenAI API feature gaps, evaluate streaming and non-streaming paths together. Custom tool calls, multi-choice Chat Completions chunks, non-text tool outputs, and similar provider payload differences must not be strict in one path and permissive or malformed in the other.
|
||||
- When a change creates new public SDK behavior, do not expose it only through hard-coded module globals. Prefer an explicit public configuration object or parameter, preserve the existing default behavior when compatibility-sensitive, and make opt-in SDK defaults explicit.
|
||||
- Append new optional fields or constructor parameters to public dataclasses and constructors. Do not insert them before existing public fields unless you also provide a compatibility layer and regression coverage for the old positional call shape.
|
||||
- Treat threshold and quota values as part of the API design when they affect runtime behavior. Distinguish OpenAI platform quota-derived values from defensive SDK defaults; if the value is not anchored in a documented platform limit, avoid making it an unconditional default-on behavior.
|
||||
- Define `None` semantics deliberately for public configuration. For example, use separate meanings for "feature disabled or no SDK limit", "use SDK default limits", and "disable only this specific limit" rather than relying on implicit truthiness checks.
|
||||
|
||||
## When to stop and confirm
|
||||
|
||||
- The change would alter behavior shipped in the latest release tag.
|
||||
- The change would modify durable external data, protocol formats, or serialized state.
|
||||
- The user explicitly asked for backward compatibility, deprecation, or migration support.
|
||||
|
||||
## Output expectations
|
||||
|
||||
When this skill materially affects the implementation approach, state the decision briefly in your reasoning or handoff, for example:
|
||||
|
||||
- `Compatibility boundary: latest release tag v0.x.y; branch-local interface rewrite, no shim needed.`
|
||||
- `Compatibility boundary: released RunState schema; preserve compatibility and add migration coverage.`
|
||||
@@ -0,0 +1,4 @@
|
||||
interface:
|
||||
display_name: "Implementation Strategy"
|
||||
short_description: "Choose a compatibility-aware implementation plan"
|
||||
default_prompt: "Use $implementation-strategy to choose the implementation approach and compatibility boundary before editing runtime code."
|
||||
@@ -0,0 +1,44 @@
|
||||
---
|
||||
name: openai-knowledge
|
||||
description: Use when working with the OpenAI API (Responses API) or OpenAI platform features (tools, streaming, Realtime API, auth, models, rate limits, MCP) and you need authoritative, up-to-date documentation (schemas, examples, limits, edge cases). Prefer the OpenAI Developer Documentation MCP server tools when available; otherwise guide the user to enable `openaiDeveloperDocs`.
|
||||
---
|
||||
|
||||
# OpenAI Knowledge
|
||||
|
||||
## Overview
|
||||
|
||||
Use the OpenAI Developer Documentation MCP server to search and fetch exact docs (markdown), then base your answer on that text instead of guessing.
|
||||
|
||||
## Workflow
|
||||
|
||||
### 1) Check whether the Docs MCP server is available
|
||||
|
||||
If the `mcp__openaiDeveloperDocs__*` tools are available, use them.
|
||||
|
||||
If you are unsure, run `codex mcp list` and check for `openaiDeveloperDocs`.
|
||||
|
||||
### 2) Use MCP tools to pull exact docs
|
||||
|
||||
- Search first, then fetch the specific page or pages.
|
||||
- `mcp__openaiDeveloperDocs__search_openai_docs` → pick the best URL.
|
||||
- `mcp__openaiDeveloperDocs__fetch_openai_doc` → retrieve the exact markdown (optionally with an `anchor`).
|
||||
- When you need endpoint schemas or parameters, use:
|
||||
- `mcp__openaiDeveloperDocs__get_openapi_spec`
|
||||
- `mcp__openaiDeveloperDocs__list_api_endpoints`
|
||||
|
||||
Base your answer on the fetched text and quote or paraphrase it precisely. Do not invent flags, field names, defaults, or limits.
|
||||
|
||||
### 3) If MCP is not configured, guide setup (do not change config unless asked)
|
||||
|
||||
Provide one of these setup options, then ask the user to restart the Codex session so the tools load:
|
||||
|
||||
- CLI:
|
||||
- `codex mcp add openaiDeveloperDocs --url https://developers.openai.com/mcp`
|
||||
- Config file (`~/.codex/config.toml`):
|
||||
- Add:
|
||||
```toml
|
||||
[mcp_servers.openaiDeveloperDocs]
|
||||
url = "https://developers.openai.com/mcp"
|
||||
```
|
||||
|
||||
Also point to: https://developers.openai.com/resources/docs-mcp#quickstart
|
||||
@@ -0,0 +1,4 @@
|
||||
interface:
|
||||
display_name: "OpenAI Knowledge"
|
||||
short_description: "Pull authoritative OpenAI platform documentation"
|
||||
default_prompt: "Use $openai-knowledge to fetch the exact OpenAI docs needed for this API or platform question."
|
||||
@@ -0,0 +1,58 @@
|
||||
---
|
||||
name: pr-draft-summary
|
||||
description: Create the required PR-ready summary block, branch suggestion, title, and draft description for openai-agents-python. Use in the final handoff after moderate-or-larger changes to runtime code, tests, examples, build/test configuration, or docs with behavior impact; skip only for trivial or conversation-only tasks, repo-meta/doc-only tasks without behavior impact, or when the user explicitly says not to include the PR draft block.
|
||||
---
|
||||
|
||||
# PR Draft Summary
|
||||
|
||||
## Purpose
|
||||
Produce the PR-ready summary required in this repository after substantive code work is complete: a concise summary plus a PR-ready title and draft description that begins with "This pull request <verb> ...". The block should be ready to paste into a PR for openai-agents-python.
|
||||
|
||||
## When to Trigger
|
||||
- The task for this repo is finished (or ready for review) and it touched runtime code, tests, examples, docs with behavior impact, or build/test configuration.
|
||||
- Treat this as the default final handoff step for substantive code work. Run it after any required verification or changeset work and before sending the "work complete" response.
|
||||
- Skip only for trivial or conversation-only tasks, repo-meta/doc-only tasks without behavior impact, or when the user explicitly says not to include the PR draft block.
|
||||
|
||||
## Inputs to Collect Automatically (do not ask the user)
|
||||
- Current branch: `git rev-parse --abbrev-ref HEAD`.
|
||||
- Working tree: `git status -sb`.
|
||||
- Untracked files: `git ls-files --others --exclude-standard` (use with `git status -sb` to ensure they are surfaced; `--stat` does not include them).
|
||||
- Changed files: `git diff --name-only` (unstaged) and `git diff --name-only --cached` (staged); sizes via `git diff --stat` and `git diff --stat --cached`.
|
||||
- Latest release tag (prefer remote-aware lookup): `LATEST_RELEASE_TAG=$(.agents/skills/final-release-review/scripts/find_latest_release_tag.sh origin 'v*' 2>/dev/null || git tag -l 'v*' --sort=-v:refname | head -n1)`.
|
||||
- Base reference (use the branch's upstream, fallback to `origin/main`):
|
||||
- `BASE_REF=$(git rev-parse --abbrev-ref --symbolic-full-name @{upstream} 2>/dev/null || echo origin/main)`.
|
||||
- `BASE_COMMIT=$(git merge-base --fork-point "$BASE_REF" HEAD || git merge-base "$BASE_REF" HEAD || echo "$BASE_REF")`.
|
||||
- Commits ahead of the base fork point: `git log --oneline --no-merges ${BASE_COMMIT}..HEAD`.
|
||||
- Category signals for this repo: runtime (`src/agents/`), tests (`tests/`), examples (`examples/`), docs (`docs/`, `mkdocs.yml`), build/test config (`pyproject.toml`, `uv.lock`, `Makefile`, `.github/`).
|
||||
|
||||
## Workflow
|
||||
1) Run the commands above without asking the user; compute `BASE_REF`/`BASE_COMMIT` first so later commands reuse them.
|
||||
2) If there are no staged/unstaged/untracked changes and no commits ahead of `${BASE_COMMIT}`, reply briefly that no code changes were detected and skip emitting the PR block.
|
||||
3) Infer change type from the touched paths listed under "Category signals"; classify as feature, fix, refactor, or docs-with-impact, and flag backward-compatibility risk only when the diff changes released public APIs, external config, persisted data, serialized state, or wire protocols. Judge that risk against `LATEST_RELEASE_TAG`, not unreleased branch-only churn.
|
||||
4) Summarize changes in 1–3 short sentences using the key paths (top 5) and `git diff --stat` output; explicitly call out untracked files from `git status -sb`/`git ls-files --others --exclude-standard` because `--stat` does not include them. If the working tree is clean but there are commits ahead of `${BASE_COMMIT}`, summarize using those commit messages.
|
||||
5) Choose the lead verb for the description: feature → `adds`, bug fix → `fixes`, refactor/perf → `improves` or `updates`, docs-only → `updates`.
|
||||
6) Suggest a branch name. If already off main, keep it; otherwise propose `feat/<slug>`, `fix/<slug>`, or `docs/<slug>` based on the primary area (e.g., `docs/pr-draft-summary-guidance`).
|
||||
7) If the current branch matches `issue-<number>` (digits only), keep that branch suggestion. Optionally pull light issue context (for example via the GitHub API) when available, but do not block or retry if it is not. When an issue number is present, reference `https://github.com/openai/openai-agents-python/issues/<number>` and include an auto-closing line such as `This pull request resolves #<number>.`.
|
||||
8) Draft the PR title and description using the template below.
|
||||
9) Output only the block in "Output Format". Keep any surrounding status note minimal and in English.
|
||||
|
||||
## Output Format
|
||||
When closing out a task, add this concise Markdown block (English only) after any brief status note unless the task falls under the documented skip cases or the user says they do not want it.
|
||||
|
||||
```
|
||||
# Pull Request Draft
|
||||
|
||||
## Branch name suggestion
|
||||
|
||||
git checkout -b <kebab-case suggestion, e.g., feat/pr-draft-summary-skill>
|
||||
|
||||
## Title
|
||||
|
||||
<single-line imperative title, which can be a commit message; if a common prefix like chore: and feat: etc., having them is preferred>
|
||||
|
||||
## Description
|
||||
|
||||
<include what you changed plus a draft pull request title and description for your local changes; start the description with prose such as "This pull request resolves/updates/adds ..." using a verb that matches the change (you can use bullets later), explain the change background (for bugs, clearly describe the bug, symptoms, or repro; for features, what is needed and why), any behavior changes or considerations to be aware of, and you do not need to mention tests you ran.>
|
||||
```
|
||||
|
||||
Keep it tight—no redundant prose around the block, and avoid repeating details between `Changes` and the description. Tests do not need to be listed unless specifically requested.
|
||||
@@ -0,0 +1,4 @@
|
||||
interface:
|
||||
display_name: "PR Draft Summary"
|
||||
short_description: "Draft the repo-ready PR title and description"
|
||||
default_prompt: "Use $pr-draft-summary to generate the PR-ready summary block, title, and draft description for the current changes."
|
||||
@@ -0,0 +1,160 @@
|
||||
---
|
||||
name: runtime-behavior-probe
|
||||
description: Plan and execute runtime-behavior investigations with temporary probe scripts, validation matrices, state controls, and findings-first reports. Use only when the user explicitly invokes this skill to verify actual runtime behavior beyond normal code-level checks, especially to uncover edge cases, undocumented behavior, or common failure modes in local or live integrations. A baseline smoke check is fine as an entry point, but do not stop at happy-path confirmation.
|
||||
---
|
||||
|
||||
# Runtime Behavior Probe
|
||||
|
||||
## Overview
|
||||
|
||||
Use this skill to investigate real runtime behavior, not to restate code or documentation. Start by planning the investigation, then execute a case matrix, record observed behavior, and report both the findings and the method used to obtain them.
|
||||
|
||||
## Core Rules
|
||||
|
||||
- Treat this skill as manual-only. Do not rely on implicit invocation.
|
||||
- A baseline success or smoke case is often the right entry point, but do not stop there when the real question involves edge cases, drift, or failure behavior.
|
||||
- Plan before running anything. Write the case matrix first, then fill it in with observed results. The matrix can live in a scratch note, a temporary file, or the probe script header.
|
||||
- Default to local or read-only probes. Consider a live service only when it is clearly relevant, then apply the lightweight gates below before you run it.
|
||||
- Size the probe to the decision. Start with the smallest matrix that can disqualify or validate the current hypothesis, then expand only when uncertainty remains.
|
||||
- Before a live probe, apply three lightweight gates:
|
||||
- Destination gate. Use only a live destination that is clearly allowed for the task.
|
||||
- Intent gate. Run the live probe only when the user explicitly wants runtime verification on that integration, or explicitly approves it after you propose the probe.
|
||||
- Data gate. If the probe will read environment variables, mutate remote state, incur material cost, or exercise non-public or user data, name the exact variable names or data class and get explicit approval first.
|
||||
- Classify each case as read-only, mutating, or costly before execution. For mutating or costly cases, or for any live case that will read environment variables, define cleanup or rollback before running the probe.
|
||||
- Use temporary files or a temporary directory for one-off probe scripts.
|
||||
- Keep temporary artifacts until the final response is drafted. Then delete them by default unless the user asked to keep them or they are needed for follow-up. Even when artifacts are deleted, keep a short run summary of the command shape, runtime context, and artifact status in the report.
|
||||
- Before executing a live probe that will read environment variables, tell the user the exact variable names you plan to use and why, then wait for explicit approval. Examples include `OPENAI_API_KEY` and other expected default names for the system under test.
|
||||
- Never print secrets, even when they come from standard environment variables that this skill may use.
|
||||
- For OpenAI API or OpenAI platform probes in this repository, use [$openai-knowledge](../openai-knowledge/SKILL.md) early to confirm contract-sensitive details such as supported parameters, field names, and limits. Use runtime probing to validate or challenge the documented behavior, not to skip the documentation pass entirely. If the docs MCP is unavailable, fall back to the official OpenAI docs and say that you used the fallback in the report.
|
||||
- For benchmark or comparison probes, make parity explicit before execution. Record what is held constant, what variable is under test, which response-shape constraints keep the comparison fair, and any usage or token counters that matter for interpreting latency or cost.
|
||||
- For OpenAI hosted tool probes, remove setup ambiguity before attributing a negative result to runtime behavior:
|
||||
- Force the tool path with the matching `tool_choice` when the question depends on tool invocation.
|
||||
- Treat `container_auto` and `container_reference` as separate cases, not interchangeable setup details.
|
||||
- Clear unsupported model or tool options first so they do not invalidate the probe.
|
||||
|
||||
## Workflow
|
||||
|
||||
1. Restate the investigation target in operational terms. Name the runtime surface, the key uncertainty, and the highest-risk behaviors to test.
|
||||
2. Do a short preflight. Check the relevant code or docs first, decide whether the question needs local or live validation, and note any repo, baseline, or release boundary that matters.
|
||||
3. Create a validation matrix before executing probes. Cover both baseline behavior and the most relevant failure or drift cases. The matrix can live in a scratch note, a temporary file, or a structured header inside the probe script.
|
||||
4. For each case, choose an execution mode up front:
|
||||
- `single-shot` for deterministic one-run checks.
|
||||
- `repeat-N` for cache, retry, streaming, interruption, rate-limit, concurrency, or other run-to-run-sensitive behavior.
|
||||
- `warm-up + repeat-N` when first-run cold-start effects could distort the result.
|
||||
Use these defaults unless the task clearly needs something else:
|
||||
- Quick screen of a repeat-sensitive question: `repeat-3`.
|
||||
- Decision-grade latency or release recommendation: `warm-up + repeat-10`.
|
||||
- Costly live cases: start at `repeat-3`, then expand only if the answer remains unclear.
|
||||
If it is genuinely unclear whether extra runs are worth the time or cost, ask the user before expanding the probe.
|
||||
5. When the question is benchmark-like or comparative, run in phases. Start with a high-signal pilot matrix against a control, then expand only the surviving candidates or unresolved cases.
|
||||
6. If the question is about a suspected regression or behavior change, add at least one known-good control case such as `origin/main`, the latest release, or the same request without the suspected option.
|
||||
7. For comparative probes, define parity before execution. Record prompt or input shape, tool-choice setup, model-settings parity, state reuse rules, and any response-shape constraint that keeps the comparison fair. If materially different output length could bias the result, record usage or token notes too.
|
||||
8. If the question asks whether one option has the same intelligence or quality as another, decide whether the matrix supports only example-pattern parity or a broader quality claim. For broader claims, add at least one harder or more open-ended case. Otherwise say explicitly that the result is limited to the covered patterns.
|
||||
9. Plan state controls before execution when hidden state could affect the result. Record whether each case uses fresh or reused state, how cache reuse or cache busting is handled, what unique IDs isolate repeated runs, and how cleanup is verified.
|
||||
10. If any live case will read environment variables, list the exact variable names and purpose for each case, then ask the user for approval before execution. Keep the approval ask short and include destination, read-only versus mutating or costly risk, exact variable names, and cleanup or rollback if relevant.
|
||||
11. Build task-specific probe scripts in a temporary location. Keep the script small, observable, and easy to discard.
|
||||
12. In `openai-agents-python`, make the runtime context explicit:
|
||||
- Run Python probes from the repository root with `uv run python` when practical.
|
||||
- Record the current commit, working directory, Python executable, and Python version.
|
||||
- Avoid accidental imports from a different checkout or site-packages location. If you must deviate from `uv run python`, say exactly why and what interpreter or environment was used instead.
|
||||
13. Execute the matrix and capture evidence. Record request shape, setup, observation summary, unexpected or negative result, error details, timing, runtime context, approved environment-variable names, repeat counts, warm-up handling, variance when relevant, cleanup behavior, and for comparisons note what was held constant plus any response-shape or usage notes that affect interpretation.
|
||||
14. Update the matrix with actual outcomes, not guesses.
|
||||
15. Keep temporary artifacts until the final response is drafted. Then delete them unless the user asked to keep them or they are needed for follow-up. Benchmark and repeat-heavy probes often need follow-up, so keeping artifacts is normal when the result may be revisited. If deleted, retain and report a short run summary.
|
||||
16. Report findings first, with unexpected or negative findings first. Then summarize how the validation was performed and which cases were covered.
|
||||
17. If the probe isolates one clear defect, you may include a short implementation hypothesis or minimal repro direction. Do not expand into a larger next-step plan unless the user asked for it.
|
||||
|
||||
## Validation Matrix
|
||||
|
||||
Use a matrix that makes the news easy to scan. Start from the runtime question and the observation summary, not just from `expected` and `pass` or `fail`.
|
||||
|
||||
Use a matrix with at least these columns:
|
||||
|
||||
- `case_id`
|
||||
- `scenario`
|
||||
- `mode`
|
||||
- `question`
|
||||
- `setup`
|
||||
- `observation_summary`
|
||||
- `result_flag`
|
||||
- `evidence`
|
||||
|
||||
Add these columns when they materially improve the investigation:
|
||||
|
||||
- `comparison_basis`
|
||||
- `variable_under_test`
|
||||
- `held_constant`
|
||||
- `output_constraint`
|
||||
- `status`
|
||||
- `confidence`
|
||||
- `state_setup`
|
||||
- `repeats`
|
||||
- `warm_up`
|
||||
- `variance`
|
||||
- `usage_note`
|
||||
- `risk_profile`
|
||||
- `env_vars`
|
||||
- `approval`
|
||||
- `control`
|
||||
|
||||
Treat `result_flag` as a fast scan field such as `unexpected`, `negative`, `expected`, or `blocked`. Use `status` only when there is a credible comparison basis, baseline, or documented contract to compare against.
|
||||
|
||||
Always consider whether the matrix should include these categories:
|
||||
|
||||
- Baseline success.
|
||||
- Control or baseline comparison when a regression is suspected.
|
||||
- Boundary input or parameter variation.
|
||||
- Invalid or unsupported input.
|
||||
- Missing or incorrect configuration.
|
||||
- Transient external failure such as timeout, network interruption, or rate limiting.
|
||||
- Retry, idempotence, or cleanup behavior.
|
||||
- Concurrency or overlapping operations when shared state or ordering may matter.
|
||||
- Open-ended quality or intelligence samples when the question is broader than pattern parity.
|
||||
|
||||
Open [validation-matrix.md](./references/validation-matrix.md) when you need a stronger prioritization model or a reusable case template.
|
||||
|
||||
## Temporary Probe Scripts
|
||||
|
||||
Write one-off scripts in a temporary file or temporary directory such as one created by `mktemp -d` or Python `tempfile`. Keep the script outside the repository by default, even when it imports code from the repository.
|
||||
|
||||
If the probe needs repository code:
|
||||
|
||||
- Run it with the repository as the working directory, or
|
||||
- Set `PYTHONPATH` or the equivalent import path explicitly.
|
||||
- In `openai-agents-python`, prefer `uv run python /tmp/probe.py` from the repository root.
|
||||
|
||||
Design the probe to maximize observability:
|
||||
|
||||
- Print or log the exact scenario being exercised.
|
||||
- Capture runtime context such as git SHA, working directory, Python executable and version, relevant package versions, model or deployment name, endpoint or base URL alias, and any retry or tool options that materially affect behavior.
|
||||
- For live probes, record only the names of environment variables that were approved for use. Never print their values.
|
||||
- Capture structured outputs when possible.
|
||||
- Preserve raw error type, message, and status code.
|
||||
- For repeat-sensitive cases, capture the attempt index, warm-up status, and any stable identifiers that help compare runs.
|
||||
- For repeated or benchmark-style probes, write both raw results and a compact summary artifact when practical.
|
||||
- Keep branching minimal so each script answers a narrow question.
|
||||
|
||||
Before deleting the temporary script or directory, keep a short run summary of the script path, command used, runtime context, and whether the evidence was kept or deleted.
|
||||
|
||||
Open [python_probe.py](./templates/python_probe.py) when you want a lightweight disposable Python probe scaffold.
|
||||
|
||||
## Reporting
|
||||
|
||||
Report in this order:
|
||||
|
||||
1. Findings. Put unexpected or negative findings first. If there was no real news, say that explicitly.
|
||||
2. Validation approach. Summarize the code used, the runtime surface exercised, the execution modes, and the case matrix coverage.
|
||||
3. Case results. Include the matrix or a condensed version of it when the case count is large.
|
||||
4. Artifact status and brief run summary. State whether temporary artifacts were deleted or kept, and provide kept paths or the retained summary.
|
||||
5. Optional implementation note. Include this only when one clear defect was isolated and a short implementation direction would help.
|
||||
|
||||
For comparative probes, the report should also say what was held constant, what variable was under test, and whether the result supports only pattern parity or a broader quality claim.
|
||||
|
||||
Open [reporting-format.md](./references/reporting-format.md) for the recommended response template.
|
||||
|
||||
## Resources
|
||||
|
||||
- Open [validation-matrix.md](./references/validation-matrix.md) to design and prioritize the case matrix.
|
||||
- Open [error-cases.md](./references/error-cases.md) to expand common failure scenarios.
|
||||
- Open [openai-runtime-patterns.md](./references/openai-runtime-patterns.md) for recurring OpenAI and Responses API probe patterns.
|
||||
- Open [reporting-format.md](./references/reporting-format.md) for the final report structure.
|
||||
- Open [python_probe.py](./templates/python_probe.py) for a minimal disposable Python probe scaffold.
|
||||
@@ -0,0 +1,6 @@
|
||||
interface:
|
||||
display_name: "Runtime Behavior Probe"
|
||||
short_description: "Plan and run runtime behavior probes"
|
||||
default_prompt: "Use $runtime-behavior-probe to investigate actual runtime behavior with a validation matrix, explicit state controls, and a findings-first report."
|
||||
policy:
|
||||
allow_implicit_invocation: false
|
||||
@@ -0,0 +1,80 @@
|
||||
# Common Error Cases
|
||||
|
||||
Use this reference to expand beyond the happy path. Favor error cases that a real user or operator is likely to hit.
|
||||
|
||||
## Configuration Errors
|
||||
|
||||
Check whether the runtime behaves differently for:
|
||||
|
||||
- Missing required environment variables.
|
||||
- Present but malformed secrets or identifiers.
|
||||
- Wrong endpoint or base URL.
|
||||
- Wrong model or deployment name.
|
||||
- Incompatible local dependency versions.
|
||||
|
||||
Look for:
|
||||
|
||||
- Error type and status code.
|
||||
- Whether the failure is immediate or delayed.
|
||||
- Whether the message is actionable.
|
||||
- Whether retrying without fixing configuration changes anything.
|
||||
|
||||
## Input Errors
|
||||
|
||||
Probe common bad-input patterns such as:
|
||||
|
||||
- Missing required fields.
|
||||
- Wrong data type.
|
||||
- Unsupported enum or option value.
|
||||
- Empty but syntactically valid input.
|
||||
- Oversized input or too many items.
|
||||
- Mutually incompatible options.
|
||||
|
||||
Prefer realistic invalid inputs over artificial nonsense. The point is to learn how the runtime fails in practice.
|
||||
|
||||
## Transport and Availability Errors
|
||||
|
||||
When networked services are involved, consider:
|
||||
|
||||
- Connection failure.
|
||||
- Read timeout.
|
||||
- Server timeout or upstream gateway error.
|
||||
- Rate limit response.
|
||||
- Partial stream interruption.
|
||||
- Reusing a connection after a failure.
|
||||
|
||||
Capture whether the client library retries automatically, whether it surfaces retry metadata, and whether the final exception preserves the original cause.
|
||||
|
||||
## State and Repetition Errors
|
||||
|
||||
Many surprising bugs appear only when an operation is repeated or interrupted:
|
||||
|
||||
- Re-submit the same request.
|
||||
- Repeat after a timeout.
|
||||
- Retry after a partial tool call or partial stream.
|
||||
- Resume after local cleanup or process restart.
|
||||
- Repeat with slightly changed inputs while reusing shared state.
|
||||
|
||||
Observe whether the operation is idempotent, duplicated, silently ignored, or left in a partial state.
|
||||
|
||||
## Concurrency Errors
|
||||
|
||||
When shared state, ordering, or isolation may matter, consider:
|
||||
|
||||
- Two overlapping requests with the same logical input.
|
||||
- Parallel runs that reuse the same cache key, session, container, or temporary resource.
|
||||
- Concurrent retries, cancellation, or cleanup racing with active work.
|
||||
- Output or event streams from one run leaking into another.
|
||||
|
||||
Capture whether the runtime serializes, rejects, duplicates, corrupts, or cross-contaminates the work.
|
||||
|
||||
## Investigation Heuristics
|
||||
|
||||
Use these heuristics to pick error cases quickly:
|
||||
|
||||
- Ask which failure a real engineer would debug first in production.
|
||||
- Ask which failure is most expensive if it is misunderstood.
|
||||
- Ask which failure would be invisible from code review alone.
|
||||
- Ask which failure path is likely to differ across environments.
|
||||
|
||||
If the error behavior is already perfectly obvious from a local validator or type system, it is usually low priority for this skill.
|
||||
@@ -0,0 +1,126 @@
|
||||
# OpenAI Runtime Patterns
|
||||
|
||||
Use this reference for recurring OpenAI investigations so you do not have to rediscover the probe strategy each time. In this repository, use [$openai-knowledge](../../openai-knowledge/SKILL.md) up front for contract-sensitive details, then use this reference to design the runtime validation. If the docs MCP is unavailable, fall back to the official OpenAI docs and say so in the report.
|
||||
|
||||
## General Rules
|
||||
|
||||
- Prefer small live probes over large harnesses.
|
||||
- Keep one script focused on one uncertainty.
|
||||
- For comparative or benchmark-like questions, start with a pilot and expand only when the answer is still unclear.
|
||||
- Capture both the request shape and the returned item types.
|
||||
- Preserve raw error payloads and status codes.
|
||||
- Record whether behavior differs between the first call and a repeated call.
|
||||
- When the question is about regression or contract drift, add a known-good control run before attributing the result to the change under investigation.
|
||||
- Keep comparison parity explicit. Record what was held constant, what variable changed, and whether output-shape or usage differences could bias the conclusion.
|
||||
- When the question depends on tool invocation, force the target path with the matching `tool_choice`.
|
||||
- Treat `container_auto` and `container_reference` as distinct setup modes, not interchangeable details.
|
||||
- Clear unsupported model or tool options before diagnosing runtime behavior.
|
||||
|
||||
## Standard Environment Variables
|
||||
|
||||
Do not read these variables automatically. Before a live probe uses any of them, tell the user the exact variable names you plan to read and why each one is needed, then wait for explicit approval. Never print their values:
|
||||
|
||||
- `OPENAI_API_KEY`
|
||||
- `OPENAI_BASE_URL`
|
||||
- `OPENAI_ORG_ID`
|
||||
- `OPENAI_PROJECT_ID`
|
||||
|
||||
If the task targets another standard integration, use that integration's expected default variable names under the same rule.
|
||||
|
||||
## Responses API Probe Patterns
|
||||
|
||||
For Responses API work, start from the uncertainty instead of from the full feature surface.
|
||||
|
||||
### Benchmark or model-switch comparisons
|
||||
|
||||
Use when you need to compare models, settings, transports, or providers with enough rigor to support a product or release decision.
|
||||
|
||||
Probe suggestions:
|
||||
|
||||
- Start with a pilot that includes one control and two or three highest-signal scenarios.
|
||||
- Keep prompt shape, tool choice, state setup, and non-tested settings aligned across candidates.
|
||||
- If the question is about speed, capture medians and, when relevant, first-token latency plus any usage note that could explain the difference.
|
||||
- If the question is about "same intelligence" or "same quality," add at least one harder or more open-ended case. Otherwise report the result as pattern parity only.
|
||||
- Expand to a larger matrix only when the pilot survives, the candidates are close, or a major runtime surface is still uncovered.
|
||||
|
||||
### Plain response behavior
|
||||
|
||||
Use when you need to confirm:
|
||||
|
||||
- The shape of returned output items.
|
||||
- Whether text appears in one item or multiple items.
|
||||
- How metadata appears in the final object.
|
||||
|
||||
Probe suggestions:
|
||||
|
||||
- Baseline call with a minimal input.
|
||||
- Same call with a slightly different instruction shape.
|
||||
- Repeat the same call to check output stability where that matters.
|
||||
|
||||
### Structured output behavior
|
||||
|
||||
Use when you need to observe:
|
||||
|
||||
- Schema rejection versus best-effort completion.
|
||||
- Handling of missing required fields.
|
||||
- Differences between model-compliant output and transport-level errors.
|
||||
|
||||
Probe suggestions:
|
||||
|
||||
- Valid schema and valid prompt.
|
||||
- Prompt likely to produce omitted fields.
|
||||
- Clearly incompatible schema or unsupported option when relevant.
|
||||
|
||||
### Tool invocation behavior
|
||||
|
||||
Use when you need to learn:
|
||||
|
||||
- When tool calls are emitted.
|
||||
- How arguments are shaped at runtime.
|
||||
- What happens when the tool fails or returns malformed output.
|
||||
|
||||
Probe suggestions:
|
||||
|
||||
- Baseline tool-call success.
|
||||
- Tool failure with a realistic exception.
|
||||
- Tool result that is syntactically valid but semantically incomplete.
|
||||
|
||||
### Hosted shell and code interpreter failure shields
|
||||
|
||||
When probing hosted tools through the Responses API, eliminate common setup ambiguity first:
|
||||
|
||||
- Force the tool path you want to test with the matching `tool_choice`. A text-only completion without forced tool choice is not a reliable negative result.
|
||||
- Treat `container_auto` and `container_reference` differently. Use `container_auto` when the probe needs fresh container provisioning or skill attachment, and use `container_reference` only to reuse existing container state.
|
||||
- Do not assume every environment field is accepted on every container mode. If the probe is about skills, validate that the chosen container mode actually supports skill attachment before treating an API error as a runtime defect.
|
||||
- Check model-specific option support before chasing unrelated failures. Unsupported reasoning or model settings can invalidate the probe before the tool path is exercised.
|
||||
- For hosted package installation, treat network-dependent setup as best-effort and separate install failures from the underlying tool behavior you are trying to observe.
|
||||
- For prompt cache investigations, keep model, instructions, tool configuration, and cache key effectively identical across repeated runs before interpreting `cached_tokens`.
|
||||
|
||||
### Streaming behavior
|
||||
|
||||
Use when the uncertainty involves:
|
||||
|
||||
- Event ordering.
|
||||
- Partial text delivery.
|
||||
- Termination after interruption.
|
||||
- Tool-call events in streams.
|
||||
|
||||
Probe suggestions:
|
||||
|
||||
- Normal streamed completion.
|
||||
- Early local cancellation.
|
||||
- Network interruption if it can be reproduced safely.
|
||||
|
||||
## What to Capture
|
||||
|
||||
For OpenAI probes, try to record:
|
||||
|
||||
- Request options that materially affect behavior.
|
||||
- Response item types and their order.
|
||||
- Whether fields are absent, null, empty, or transformed.
|
||||
- Server status and error payload details for failures.
|
||||
- Retry and backoff hints when present.
|
||||
- Stable identifiers that help compare repeated runs, such as request IDs, response IDs, tool call IDs, or container IDs when available.
|
||||
- Which environment-variable names were approved for the probe when live credentials were required.
|
||||
|
||||
Do not spend time rediscovering static documentation unless the runtime result seems to contradict what you expected. The value of this skill is in the observed behavior.
|
||||
@@ -0,0 +1,118 @@
|
||||
# Reporting Format
|
||||
|
||||
Lead with findings, not process. The user asked for investigation results, so the answer should start with the most important observed behaviors. Put the real news first.
|
||||
|
||||
## Recommended Order
|
||||
|
||||
1. Findings.
|
||||
2. Validation approach.
|
||||
3. Case matrix or condensed case summary.
|
||||
4. Artifact status and brief run summary.
|
||||
5. Optional implementation note.
|
||||
|
||||
## Findings Section
|
||||
|
||||
Make each finding answer one user-relevant question. Good findings usually include:
|
||||
|
||||
- What was observed.
|
||||
- Why it matters.
|
||||
- The condition under which it happens.
|
||||
- What was held constant when the finding comes from a comparison probe.
|
||||
- `scope`: The boundary of the finding, such as commit, model, Python version, live vs local, or repeat mode.
|
||||
- `confidence`: `high`, `medium`, or `low`.
|
||||
|
||||
Avoid burying the main result under setup details.
|
||||
|
||||
Put `unexpected` or `negative` findings first. If there were no unexpected or negative findings in the executed cases, say that explicitly before the rest of the findings section.
|
||||
|
||||
If the probe was comparative, say whether the result supports:
|
||||
|
||||
- Pattern parity only.
|
||||
- A broader quality claim.
|
||||
|
||||
Do not imply a broader quality equivalence than the executed cases justify.
|
||||
|
||||
## Validation Approach Section
|
||||
|
||||
Summarize:
|
||||
|
||||
- The runtime surface you exercised.
|
||||
- The shape of the probe code, in overview only.
|
||||
- Which categories of cases you covered.
|
||||
- Which execution modes you used, including repeat counts or warm-up handling when relevant.
|
||||
- Whether live credentials or external services were used.
|
||||
- Any important state controls such as fresh state, cache reuse, cache busting, unique IDs, or cleanup verification.
|
||||
- For comparison probes, what was held constant, what was varied, and whether output-shape or usage differences could still influence the conclusion.
|
||||
- Whether the usual docs path or an official-docs fallback was used for contract-sensitive checks.
|
||||
|
||||
Keep this concise. The user needs enough detail to trust the result, not a line-by-line replay of the script.
|
||||
|
||||
## Case Summary
|
||||
|
||||
Include either the full matrix or a condensed summary. At minimum, show:
|
||||
|
||||
- Which scenarios were executed.
|
||||
- Whether the run was a quick pilot, an expanded matrix, or both.
|
||||
- Which ones produced `unexpected` or `negative` results.
|
||||
- Which ones passed or failed when a real comparison basis existed.
|
||||
- Which cases were blocked.
|
||||
- Where the supporting evidence lived, or that it was deleted.
|
||||
|
||||
If the matrix is large, show the highest-value cases in the main response and keep the rest as a compact appendix or note.
|
||||
|
||||
## Artifact Status And Brief Run Summary
|
||||
|
||||
State one of these explicitly:
|
||||
|
||||
- Temporary artifacts were kept until the final response was drafted, then deleted after validation.
|
||||
- Temporary artifacts were kept at `<path>` because the user asked to keep them.
|
||||
- Temporary artifacts were kept at `<path>` because they are needed for follow-up analysis.
|
||||
|
||||
Even if artifacts were deleted, retain a short run summary such as:
|
||||
|
||||
- Probe command or runner shape.
|
||||
- Runtime context summary such as commit, Python executable, Python version, or model.
|
||||
- Artifact path and final status.
|
||||
|
||||
For benchmark or repeat-heavy probes, keeping artifacts for follow-up is often the right default even when the immediate report is done.
|
||||
|
||||
## Optional Implementation Note
|
||||
|
||||
Include this only when one clear defect was isolated and a short implementation hypothesis or minimal repro direction would help. Keep it brief. Do not turn the report into a broader next-step plan unless the user asked for that.
|
||||
|
||||
## Compact Template
|
||||
|
||||
Use this outline when you need a fast structure:
|
||||
|
||||
Findings:
|
||||
- <finding 1>
|
||||
held constant: <prompt/tool/state settings kept the same, if comparative>
|
||||
scope: <commit/model/python/live-local/repeat-mode>
|
||||
confidence: <high|medium|low>
|
||||
- <finding 2>
|
||||
held constant: <prompt/tool/state settings kept the same, if comparative>
|
||||
scope: <commit/model/python/live-local/repeat-mode>
|
||||
confidence: <high|medium|low>
|
||||
|
||||
Validation approach:
|
||||
- Surface: <what was exercised>
|
||||
- Probe code: <brief overview>
|
||||
- Coverage: <success, edge, error, repeat-sensitive, and quality categories>
|
||||
- Execution modes: <single-shot|repeat-N|warm-up + repeat-N>
|
||||
- Comparison parity: <what was held constant and what varied, if comparative>
|
||||
- Docs source: <MCP or official-docs fallback, if relevant>
|
||||
|
||||
Case summary:
|
||||
| case_id | scenario | result_flag | status | note |
|
||||
| --- | --- | --- | --- | --- |
|
||||
| S1 | ... | expected | pass | ... |
|
||||
| E1 | ... | negative | fail | ... |
|
||||
|
||||
Artifact status and brief run summary:
|
||||
- Temporary artifacts were kept until the final response was drafted, then deleted.
|
||||
- Summary: <command/runtime-context/artifact-status summary>
|
||||
|
||||
Optional implementation note:
|
||||
- <brief hypothesis or minimal repro direction>
|
||||
|
||||
Adjust the format to the task, but preserve the ordering.
|
||||
@@ -0,0 +1,137 @@
|
||||
# Validation Matrix
|
||||
|
||||
Use the matrix to decide what to probe before writing scripts. The goal is not exhaustive combinatorics; the goal is high-value coverage that is visible, explainable, and likely to reveal runtime surprises. The matrix should make the real news easy to scan.
|
||||
|
||||
## Minimum Columns
|
||||
|
||||
Use these columns unless the task clearly needs more:
|
||||
|
||||
- `case_id`: Stable identifier such as `S1`, `E3`, or `R2`.
|
||||
- `scenario`: Short description of the behavior under test.
|
||||
- `mode`: `single-shot`, `repeat-N`, or `warm-up + repeat-N`.
|
||||
- `question`: The concrete runtime uncertainty this case is answering.
|
||||
- `setup`: Inputs, environment, or preconditions required for the case.
|
||||
- `observation_summary`: A compact summary of what actually happened.
|
||||
- `result_flag`: `unexpected`, `negative`, `expected`, or `blocked`.
|
||||
- `evidence`: Path, log reference, or `deleted`.
|
||||
|
||||
Add these columns when they materially improve the investigation:
|
||||
|
||||
- `comparison_basis`: The baseline, docs, or prior behavior you are comparing against.
|
||||
- `variable_under_test`: The single factor that is intentionally changing in a comparison case.
|
||||
- `held_constant`: Prompt shape, tool setup, model settings, or state rules that were intentionally kept the same.
|
||||
- `output_constraint`: Any schema, length, or response-shape constraint used to keep the comparison fair.
|
||||
- `status`: Use `pass`, `fail`, `unexpected-pass`, `unexpected-fail`, or `blocked` only when there is a credible comparison basis or control.
|
||||
- `confidence`: `high`, `medium`, or `low`.
|
||||
- `state_setup`: Fresh or reused state, cache strategy, unique IDs, and cleanup checks.
|
||||
- `repeats`: Number of measured runs.
|
||||
- `warm_up`: Whether a warm-up run was used and why.
|
||||
- `variance`: Any useful spread or instability note across repeated runs.
|
||||
- `usage_note`: Token, usage, or output-length note when it materially affects interpretation.
|
||||
- `control`: Known-good comparison point for regression or behavior-change questions.
|
||||
- `risk_profile`: `read-only`, `mutating`, or `costly` for live probes.
|
||||
- `env_vars`: Exact environment-variable names the case plans to read.
|
||||
- `approval`: `not-needed`, `pending`, or `approved` for cases that need user permission before execution.
|
||||
|
||||
Use `result_flag` as the fast scan field. It is what makes unexpected or negative findings jump out before the reader studies the full report.
|
||||
|
||||
Use `status` only when you have a real comparison basis. If the case is exploratory and there is no trustworthy baseline, prefer a strong `observation_summary` plus `result_flag` and `confidence` instead of pretending the result is a clean pass or fail.
|
||||
|
||||
## Choosing Execution Mode
|
||||
|
||||
Pick an execution mode before you run the case:
|
||||
|
||||
- Use `single-shot` for deterministic, one-run checks.
|
||||
- Use `repeat-N` automatically when the question involves cache behavior, retries, streaming, interruptions, rate limiting, concurrency, or other run-to-run-sensitive behavior.
|
||||
- Use `warm-up + repeat-N` when the first run is likely to include cold-start effects such as container provisioning, import caches, or prompt-cache population.
|
||||
|
||||
Use these defaults unless the task clearly needs something else:
|
||||
|
||||
- `repeat-3` for a quick screen of a repeat-sensitive question.
|
||||
- `warm-up + repeat-10` for decision-grade latency comparisons or release-facing recommendations.
|
||||
- For costly live probes, start at `repeat-3` and expand only if the answer is still unclear.
|
||||
|
||||
If it is genuinely unclear whether extra runs are worth the time or cost, ask the user before expanding the probe.
|
||||
|
||||
## Phase The Matrix
|
||||
|
||||
When the question is comparative or benchmark-like, do not jump straight to the largest matrix.
|
||||
|
||||
Start with a pilot:
|
||||
|
||||
- One control.
|
||||
- One or two highest-signal success cases.
|
||||
- The smallest repeat count that can disqualify a weak candidate quickly.
|
||||
|
||||
Expand only when:
|
||||
|
||||
- The candidate survives the pilot.
|
||||
- The results are close enough that more samples matter.
|
||||
- A major runtime surface is still uncovered.
|
||||
- The user explicitly wants decision-grade evidence.
|
||||
|
||||
## Coverage Categories
|
||||
|
||||
Try to cover at least one case from each relevant category:
|
||||
|
||||
- `success`: Normal behavior that should work.
|
||||
- `control`: Known-good comparison such as `origin/main`, the latest release, or the same request without the suspected option.
|
||||
- `boundary`: Size, count, or parameter limits near a plausible edge.
|
||||
- `invalid`: Bad inputs or unsupported combinations.
|
||||
- `misconfig`: Missing key, wrong endpoint, bad permissions, or incompatible local setup.
|
||||
- `transient`: Timeout, temporary server issue, network breakage, or rate limiting.
|
||||
- `recovery`: Retry behavior, partial completion, duplicate submission, or cleanup.
|
||||
- `concurrency`: Overlapping operations when shared state, ordering, or isolation may matter.
|
||||
- `quality`: A harder or more open-ended sample when the user is asking about model intelligence, not just workflow parity.
|
||||
|
||||
If time is limited, prioritize categories in this order:
|
||||
|
||||
1. Known-good control when the question implies regression or drift.
|
||||
2. Highest-risk success case.
|
||||
3. Most plausible user-facing failure.
|
||||
4. Most likely edge case with ambiguous behavior.
|
||||
5. Cleanup or retry semantics.
|
||||
6. Lower-probability extremes.
|
||||
|
||||
## Matrix Template
|
||||
|
||||
Use this compact template:
|
||||
|
||||
| case_id | scenario | mode | question | setup | state_setup | variable_under_test | held_constant | comparison_basis | observation_summary | result_flag | status | evidence |
|
||||
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
|
||||
| K1 | Known-good control | single-shot | Does the baseline still show the expected behavior? | Same probe against baseline target | Fresh state | none | current probe shape | `origin/main` or latest release | pending | pending | pending | pending |
|
||||
| S1 | Baseline success | single-shot | What does the normal success path look like at runtime? | Valid config and representative input | Fresh state | none | representative input and setup | current docs or local expectation | pending | pending | pending | pending |
|
||||
| R1 | Cache or retry behavior | warm-up + repeat-N | Does behavior change after the first run or across retries? | Same request repeated under controlled settings | Cache key or retry setup recorded | reuse versus fresh state | prompt shape and tool setup | same request without reuse, or docs if available | pending | pending | pending | pending |
|
||||
| C1 | Model comparison pilot | warm-up + repeat-N | Does candidate B preserve the covered behavior while improving latency? | Same scenario across two models | Fresh state and stable IDs | model name | prompt shape, tool choice, and model settings parity | control model in the same probe | pending | pending | pending | pending |
|
||||
| E1 | Invalid input | single-shot | How does the runtime reject a realistic bad input? | Missing required field | Fresh state | invalid field value | same request with valid field | same request with valid field | pending | pending | pending | pending |
|
||||
| X1 | Concurrent overlap | repeat-N | Do overlapping runs interfere with each other? | Two or more overlapping operations | Unique IDs plus cleanup verification | overlap timing | same logical input | same request serialized, if available | pending | pending | pending | pending |
|
||||
|
||||
## Recording Results
|
||||
|
||||
Keep `question` unchanged after execution. Put the actual behavior in `observation_summary`, then mark the scan-friendly `result_flag`.
|
||||
|
||||
Use these `result_flag` values consistently:
|
||||
|
||||
- `unexpected`: The result diverged from the best current understanding in a surprising way.
|
||||
- `negative`: The result exposed a user-relevant failure, risk, or sharp edge.
|
||||
- `expected`: The result matched the current understanding and did not reveal new risk.
|
||||
- `blocked`: The case did not produce a trustworthy observation.
|
||||
|
||||
Only fill `status` when there is a credible comparison basis. Otherwise use `observation_summary`, `result_flag`, and `confidence` to communicate what was learned without over-claiming certainty.
|
||||
|
||||
For comparison cases, use `observation_summary` and the final report to say whether the evidence supports pattern parity only or a broader quality claim.
|
||||
|
||||
If a case reveals a new branch of behavior, add a follow-up case instead of overloading the original one.
|
||||
|
||||
## Evidence Discipline
|
||||
|
||||
Treat a case as incomplete when:
|
||||
|
||||
- The observed output omits the key result you were testing.
|
||||
- The script mixed multiple questions and the result is ambiguous.
|
||||
- Hidden state, cache behavior, or previous runs may have influenced the result and were not controlled or documented.
|
||||
- The question is whether behavior changed, but the case has no credible control or baseline to compare against.
|
||||
- The case plans to read environment variables, but the exact variable names were not approved by the user before execution.
|
||||
- The case was repeat-sensitive, but it ran only once without a clear rationale.
|
||||
|
||||
When this happens, narrow the probe and rerun. A smaller script with a cleaner result is better than a more complicated script that is hard to trust.
|
||||
@@ -0,0 +1,227 @@
|
||||
"""Disposable Python probe scaffold.
|
||||
|
||||
Copy this file to a temporary location and adapt it for one narrow question.
|
||||
Recommended usage from the repository root:
|
||||
|
||||
uv run python /tmp/probe.py
|
||||
|
||||
If you want structured artifacts for repeat-heavy or benchmark probes:
|
||||
|
||||
PROBE_OUTPUT_DIR=/tmp/probe-run uv run python /tmp/probe.py
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import os
|
||||
import platform
|
||||
import shutil
|
||||
import statistics
|
||||
import subprocess
|
||||
import sys
|
||||
import time
|
||||
import uuid
|
||||
from collections import Counter, defaultdict
|
||||
from importlib import metadata
|
||||
from pathlib import Path
|
||||
|
||||
SCENARIO = "replace-me"
|
||||
RUN_LABEL = "replace-me"
|
||||
MODE = "single-shot"
|
||||
APPROVED_ENV_VARS: list[str] = []
|
||||
OUTPUT_DIR_ENV = "PROBE_OUTPUT_DIR"
|
||||
|
||||
RESULTS: list[dict[str, object]] = []
|
||||
|
||||
|
||||
def _git_value(*args: str) -> str:
|
||||
result = subprocess.run(
|
||||
["git", *args],
|
||||
check=False,
|
||||
capture_output=True,
|
||||
text=True,
|
||||
)
|
||||
if result.returncode != 0:
|
||||
return "unknown"
|
||||
return result.stdout.strip() or "unknown"
|
||||
|
||||
|
||||
def _package_version(name: str) -> str | None:
|
||||
try:
|
||||
return metadata.version(name)
|
||||
except metadata.PackageNotFoundError:
|
||||
return None
|
||||
|
||||
|
||||
def _output_dir() -> Path | None:
|
||||
value = os.getenv(OUTPUT_DIR_ENV)
|
||||
if not value:
|
||||
return None
|
||||
return Path(value)
|
||||
|
||||
|
||||
def _write_json(path: Path, payload: object) -> None:
|
||||
path.parent.mkdir(parents=True, exist_ok=True)
|
||||
path.write_text(json.dumps(payload, indent=2, sort_keys=True) + "\n")
|
||||
|
||||
|
||||
def emit(kind: str, **payload: object) -> None:
|
||||
print(
|
||||
json.dumps(
|
||||
{
|
||||
"ts": round(time.time(), 3),
|
||||
"kind": kind,
|
||||
**payload,
|
||||
},
|
||||
sort_keys=True,
|
||||
)
|
||||
)
|
||||
|
||||
|
||||
def runtime_context() -> dict[str, object]:
|
||||
approved = {name: ("set" if os.getenv(name) else "unset") for name in APPROVED_ENV_VARS}
|
||||
package_versions = {
|
||||
name: version
|
||||
for name in ("openai", "agents")
|
||||
if (version := _package_version(name)) is not None
|
||||
}
|
||||
return {
|
||||
"scenario": SCENARIO,
|
||||
"run_label": RUN_LABEL,
|
||||
"mode": MODE,
|
||||
"cwd": os.getcwd(),
|
||||
"script_path": str(Path(__file__).resolve()),
|
||||
"python_executable": sys.executable,
|
||||
"python_version": sys.version.split()[0],
|
||||
"platform": platform.platform(),
|
||||
"git_commit": _git_value("rev-parse", "HEAD"),
|
||||
"git_branch": _git_value("rev-parse", "--abbrev-ref", "HEAD"),
|
||||
"uv_path": shutil.which("uv"),
|
||||
"package_versions": package_versions,
|
||||
"approved_env_vars": approved,
|
||||
"output_dir": str(_output_dir()) if _output_dir() else None,
|
||||
}
|
||||
|
||||
|
||||
def start_case(case_id: str, *, mode: str = MODE, note: str | None = None) -> None:
|
||||
emit("case_start", case_id=case_id, mode=mode, note=note)
|
||||
|
||||
|
||||
def record_case_result(
|
||||
case_id: str,
|
||||
observation_summary: str,
|
||||
result_flag: str,
|
||||
*,
|
||||
mode: str = MODE,
|
||||
is_warmup: bool = False,
|
||||
total_latency_s: float | None = None,
|
||||
first_token_latency_s: float | None = None,
|
||||
metrics: dict[str, object] | None = None,
|
||||
error: str | None = None,
|
||||
) -> None:
|
||||
payload: dict[str, object] = {
|
||||
"case_id": case_id,
|
||||
"mode": mode,
|
||||
"is_warmup": is_warmup,
|
||||
"observation_summary": observation_summary,
|
||||
"result_flag": result_flag,
|
||||
"metrics": metrics or {},
|
||||
"error": error,
|
||||
}
|
||||
if total_latency_s is not None:
|
||||
payload["total_latency_s"] = total_latency_s
|
||||
if first_token_latency_s is not None:
|
||||
payload["first_token_latency_s"] = first_token_latency_s
|
||||
RESULTS.append(payload)
|
||||
emit("case_result", **payload)
|
||||
|
||||
|
||||
def summarize_results() -> dict[str, object]:
|
||||
by_case: defaultdict[str, list[dict[str, object]]] = defaultdict(list)
|
||||
for result in RESULTS:
|
||||
by_case[str(result["case_id"])].append(result)
|
||||
|
||||
summary_cases: dict[str, object] = {}
|
||||
for case_id, items in by_case.items():
|
||||
measured = [item for item in items if not bool(item.get("is_warmup"))]
|
||||
latencies = [
|
||||
float(item["total_latency_s"])
|
||||
for item in measured
|
||||
if item.get("total_latency_s") is not None
|
||||
]
|
||||
first_token_latencies = [
|
||||
float(item["first_token_latency_s"])
|
||||
for item in measured
|
||||
if item.get("first_token_latency_s") is not None
|
||||
]
|
||||
result_flags = Counter(str(item["result_flag"]) for item in measured or items)
|
||||
observations = [str(item["observation_summary"]) for item in (measured or items)[:3]]
|
||||
summary_cases[case_id] = {
|
||||
"mode": str(items[-1]["mode"]),
|
||||
"runs": len(measured),
|
||||
"warmups": len(items) - len(measured),
|
||||
"result_flags": dict(result_flags),
|
||||
"median_total_latency_s": (statistics.median(latencies) if latencies else None),
|
||||
"mean_total_latency_s": statistics.mean(latencies) if latencies else None,
|
||||
"median_first_token_latency_s": (
|
||||
statistics.median(first_token_latencies) if first_token_latencies else None
|
||||
),
|
||||
"observations": observations,
|
||||
}
|
||||
|
||||
return {
|
||||
"scenario": SCENARIO,
|
||||
"run_label": RUN_LABEL,
|
||||
"mode": MODE,
|
||||
"result_count": len(RESULTS),
|
||||
"cases": summary_cases,
|
||||
"result_flags": dict(Counter(str(item["result_flag"]) for item in RESULTS)),
|
||||
}
|
||||
|
||||
|
||||
def finalize(exit_code: int) -> None:
|
||||
metadata_payload = {
|
||||
"exit_code": exit_code,
|
||||
"runtime_context": runtime_context(),
|
||||
}
|
||||
summary_payload = summarize_results()
|
||||
emit("summary", metadata=metadata_payload, summary=summary_payload)
|
||||
|
||||
output_dir = _output_dir()
|
||||
if not output_dir:
|
||||
return
|
||||
|
||||
metadata_path = output_dir / "metadata.json"
|
||||
results_path = output_dir / "results.json"
|
||||
summary_path = output_dir / "summary.json"
|
||||
_write_json(metadata_path, metadata_payload)
|
||||
_write_json(results_path, RESULTS)
|
||||
_write_json(summary_path, summary_payload)
|
||||
emit(
|
||||
"artifact_paths",
|
||||
metadata_path=str(metadata_path),
|
||||
results_path=str(results_path),
|
||||
summary_path=str(summary_path),
|
||||
)
|
||||
|
||||
|
||||
def main() -> int:
|
||||
case_id = os.getenv("PROBE_CASE_ID", f"case-{uuid.uuid4().hex[:8]}")
|
||||
emit("banner", context=runtime_context())
|
||||
start_case(case_id)
|
||||
|
||||
# Replace this block with the narrow runtime question you want to test.
|
||||
observation = "replace-me"
|
||||
result_flag = "expected"
|
||||
|
||||
record_case_result(
|
||||
case_id=case_id,
|
||||
observation_summary=observation,
|
||||
result_flag=result_flag,
|
||||
)
|
||||
finalize(exit_code=0)
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -0,0 +1,42 @@
|
||||
---
|
||||
name: test-coverage-improver
|
||||
description: 'Improve test coverage in the OpenAI Agents Python repository: run `make coverage`, inspect coverage artifacts, identify low-coverage files, propose high-impact tests, and confirm with the user before writing tests.'
|
||||
---
|
||||
|
||||
# Test Coverage Improver
|
||||
|
||||
## Overview
|
||||
|
||||
Use this skill whenever coverage needs assessment or improvement (coverage regressions, failing thresholds, or user requests for stronger tests). It runs the coverage suite, analyzes results, highlights the biggest gaps, and prepares test additions while confirming with the user before changing code.
|
||||
|
||||
## Quick Start
|
||||
|
||||
1. From the repo root run `make coverage` to regenerate `.coverage` data and `coverage.xml`.
|
||||
2. Collect artifacts: `.coverage` and `coverage.xml`, plus the console output from `coverage report -m` for drill-downs.
|
||||
3. Summarize coverage: total percentages, lowest files, and uncovered lines/paths.
|
||||
4. Draft test ideas per file: scenario, behavior under test, expected outcome, and likely coverage gain.
|
||||
5. Ask the user for approval to implement the proposed tests; pause until they agree.
|
||||
6. After approval, write the tests in `tests/`, rerun `make coverage`, and then run `$code-change-verification` before marking work complete.
|
||||
|
||||
## Workflow Details
|
||||
|
||||
- **Run coverage**: Execute `make coverage` at repo root. Avoid watch flags and keep prior coverage artifacts only if comparing trends.
|
||||
- **Parse summaries efficiently**:
|
||||
- Prefer the console output from `coverage report -m` for file-level totals; fallback to `coverage.xml` for tooling or spreadsheets.
|
||||
- Use `uv run coverage html` to generate `htmlcov/index.html` if you need an interactive drill-down.
|
||||
- **Prioritize targets**:
|
||||
- Public APIs or shared utilities in `src/agents/` before examples or docs.
|
||||
- Files with low statement coverage or newly added code at 0%.
|
||||
- Recent bug fixes or risky code paths (error handling, retries, timeouts, concurrency).
|
||||
- **Design impactful tests**:
|
||||
- Hit uncovered paths: error cases, boundary inputs, optional flags, and cancellation/timeouts.
|
||||
- Cover combinational logic rather than trivial happy paths.
|
||||
- Place tests under `tests/` and avoid flaky async timing.
|
||||
- **Coordinate with the user**: Present a numbered, concise list of proposed test additions and expected coverage gains. Ask explicitly before editing code or fixtures.
|
||||
- **After implementation**: Rerun coverage, report the updated summary, and note any remaining low-coverage areas.
|
||||
|
||||
## Notes
|
||||
|
||||
- Keep any added comments or code in English.
|
||||
- Do not create `scripts/`, `references/`, or `assets/` unless needed later.
|
||||
- If coverage artifacts are missing or stale, rerun `pnpm test:coverage` instead of guessing.
|
||||
@@ -0,0 +1,4 @@
|
||||
interface:
|
||||
display_name: "Test Coverage Improver"
|
||||
short_description: "Analyze coverage gaps and propose high-impact tests"
|
||||
default_prompt: "Use $test-coverage-improver to analyze coverage gaps, propose high-impact tests, and update coverage after approval."
|
||||
@@ -0,0 +1,4 @@
|
||||
#:schema https://developers.openai.com/codex/config-schema.json
|
||||
|
||||
[features]
|
||||
codex_hooks = true
|
||||
@@ -0,0 +1,15 @@
|
||||
{
|
||||
"hooks": {
|
||||
"Stop": [
|
||||
{
|
||||
"hooks": [
|
||||
{
|
||||
"type": "command",
|
||||
"command": "uv run python \"$(git rev-parse --show-toplevel)/.codex/hooks/stop_repo_tidy.py\"",
|
||||
"timeout": 20
|
||||
}
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,222 @@
|
||||
#!/usr/bin/env python3
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import hashlib
|
||||
import json
|
||||
import subprocess
|
||||
import sys
|
||||
import tempfile
|
||||
from dataclasses import asdict, dataclass
|
||||
from pathlib import Path
|
||||
|
||||
MAX_RUFF_FIX_FILES = 20
|
||||
PYTHON_SUFFIXES = {".py", ".pyi"}
|
||||
|
||||
|
||||
@dataclass
|
||||
class HookState:
|
||||
last_tidy_fingerprint: str | None = None
|
||||
|
||||
|
||||
def write_stop_block(reason: str, system_message: str) -> None:
|
||||
sys.stdout.write(
|
||||
json.dumps(
|
||||
{
|
||||
"decision": "block",
|
||||
"reason": reason,
|
||||
"systemMessage": system_message,
|
||||
}
|
||||
)
|
||||
)
|
||||
|
||||
|
||||
def run_command(cwd: str, *args: str) -> subprocess.CompletedProcess[str]:
|
||||
try:
|
||||
return subprocess.run(
|
||||
args,
|
||||
cwd=cwd,
|
||||
capture_output=True,
|
||||
check=False,
|
||||
text=True,
|
||||
)
|
||||
except FileNotFoundError as exc:
|
||||
return subprocess.CompletedProcess(args, returncode=127, stdout="", stderr=str(exc))
|
||||
|
||||
|
||||
def run_git(cwd: str, *args: str) -> subprocess.CompletedProcess[str]:
|
||||
return run_command(cwd, "git", *args)
|
||||
|
||||
|
||||
def git_root(cwd: str) -> str:
|
||||
result = run_git(cwd, "rev-parse", "--show-toplevel")
|
||||
if result.returncode != 0:
|
||||
raise RuntimeError(result.stderr.strip() or "git root lookup failed")
|
||||
return result.stdout.strip()
|
||||
|
||||
|
||||
def parse_status_paths(repo_root: str) -> list[str]:
|
||||
unstaged = run_git(repo_root, "diff", "--name-only", "--diff-filter=ACMR")
|
||||
untracked = run_git(repo_root, "ls-files", "--others", "--exclude-standard")
|
||||
if unstaged.returncode != 0 or untracked.returncode != 0:
|
||||
return []
|
||||
|
||||
paths = {
|
||||
line.strip()
|
||||
for result in (unstaged, untracked)
|
||||
for line in result.stdout.splitlines()
|
||||
if line.strip()
|
||||
}
|
||||
return sorted(paths)
|
||||
|
||||
|
||||
def untracked_paths(repo_root: str, paths: list[str]) -> set[str]:
|
||||
if not paths:
|
||||
return set()
|
||||
|
||||
result = run_git(repo_root, "ls-files", "--others", "--exclude-standard", "--", *paths)
|
||||
if result.returncode != 0:
|
||||
return set()
|
||||
|
||||
return {line.strip() for line in result.stdout.splitlines() if line.strip()}
|
||||
|
||||
|
||||
def fingerprint_for_paths(repo_root: str, paths: list[str]) -> str | None:
|
||||
if not paths:
|
||||
return None
|
||||
|
||||
repo_root_path = Path(repo_root)
|
||||
untracked = untracked_paths(repo_root, paths)
|
||||
tracked_paths = [file_path for file_path in paths if file_path not in untracked]
|
||||
diff_parts: list[str] = []
|
||||
|
||||
if tracked_paths:
|
||||
diff = run_git(repo_root, "diff", "--no-ext-diff", "--binary", "--", *tracked_paths)
|
||||
if diff.returncode == 0:
|
||||
diff_parts.append(diff.stdout)
|
||||
|
||||
for file_path in sorted(untracked):
|
||||
try:
|
||||
digest = hashlib.sha256((repo_root_path / file_path).read_bytes()).hexdigest()
|
||||
except OSError:
|
||||
continue
|
||||
diff_parts.append(f"untracked:{file_path}:{digest}")
|
||||
|
||||
if not diff_parts:
|
||||
return None
|
||||
|
||||
return hashlib.sha256("\n".join(diff_parts).encode("utf-8")).hexdigest()
|
||||
|
||||
|
||||
def state_dir() -> Path:
|
||||
return Path(tempfile.gettempdir()) / "openai-agents-python-codex-hooks"
|
||||
|
||||
|
||||
def state_path(session_id: str, repo_root: str) -> Path:
|
||||
root_hash = hashlib.sha256(repo_root.encode("utf-8")).hexdigest()[:12]
|
||||
safe_session_id = "".join(
|
||||
ch if ch.isascii() and (ch.isalnum() or ch in "._-") else "_" for ch in session_id
|
||||
)
|
||||
return state_dir() / f"{safe_session_id}-{root_hash}.json"
|
||||
|
||||
|
||||
def load_state(session_id: str, repo_root: str) -> HookState:
|
||||
file_path = state_path(session_id, repo_root)
|
||||
if not file_path.exists():
|
||||
return HookState()
|
||||
|
||||
try:
|
||||
payload = json.loads(file_path.read_text())
|
||||
except (OSError, json.JSONDecodeError):
|
||||
return HookState()
|
||||
|
||||
return HookState(last_tidy_fingerprint=payload.get("last_tidy_fingerprint"))
|
||||
|
||||
|
||||
def save_state(session_id: str, repo_root: str, state: HookState) -> None:
|
||||
file_path = state_path(session_id, repo_root)
|
||||
file_path.parent.mkdir(parents=True, exist_ok=True)
|
||||
file_path.write_text(json.dumps(asdict(state), indent=2))
|
||||
|
||||
|
||||
def lint_fix_paths(repo_root: str) -> list[str]:
|
||||
return [
|
||||
file_path
|
||||
for file_path in parse_status_paths(repo_root)
|
||||
if Path(file_path).suffix in PYTHON_SUFFIXES
|
||||
]
|
||||
|
||||
|
||||
def main() -> None:
|
||||
try:
|
||||
payload = json.loads(sys.stdin.read() or "null")
|
||||
except json.JSONDecodeError:
|
||||
return
|
||||
|
||||
if not isinstance(payload, dict):
|
||||
return
|
||||
|
||||
session_id = payload.get("session_id")
|
||||
cwd = payload.get("cwd")
|
||||
if not isinstance(session_id, str) or not isinstance(cwd, str):
|
||||
return
|
||||
|
||||
if payload.get("stop_hook_active"):
|
||||
return
|
||||
|
||||
repo_root = git_root(cwd)
|
||||
current_paths = lint_fix_paths(repo_root)
|
||||
if not current_paths or len(current_paths) > MAX_RUFF_FIX_FILES:
|
||||
return
|
||||
|
||||
state = load_state(session_id, repo_root)
|
||||
current_fingerprint = fingerprint_for_paths(repo_root, current_paths)
|
||||
if current_fingerprint is None or state.last_tidy_fingerprint == current_fingerprint:
|
||||
return
|
||||
|
||||
format_result = run_command(repo_root, "uv", "run", "ruff", "format", "--", *current_paths)
|
||||
check_result: subprocess.CompletedProcess[str] | None = None
|
||||
if format_result.returncode == 0:
|
||||
check_result = run_command(
|
||||
repo_root,
|
||||
"uv",
|
||||
"run",
|
||||
"ruff",
|
||||
"check",
|
||||
"--fix",
|
||||
"--",
|
||||
*current_paths,
|
||||
)
|
||||
|
||||
if format_result.returncode != 0:
|
||||
write_stop_block(
|
||||
"`uv run ruff format -- ...` failed for the touched Python files. "
|
||||
"Review the formatting step before wrapping up.",
|
||||
"Repo hook: targeted Ruff format failed.",
|
||||
)
|
||||
return
|
||||
|
||||
if check_result and check_result.returncode != 0:
|
||||
write_stop_block(
|
||||
"`uv run ruff check --fix -- ...` failed for the touched Python files. "
|
||||
"Review the lint output before wrapping up.",
|
||||
"Repo hook: targeted Ruff lint fix failed.",
|
||||
)
|
||||
return
|
||||
|
||||
updated_paths = lint_fix_paths(repo_root)
|
||||
updated_fingerprint = fingerprint_for_paths(repo_root, updated_paths)
|
||||
state.last_tidy_fingerprint = updated_fingerprint
|
||||
save_state(session_id, repo_root, state)
|
||||
|
||||
if updated_fingerprint != current_fingerprint:
|
||||
write_stop_block(
|
||||
"I ran targeted tidy steps on the touched Python files "
|
||||
"(`ruff format` and `ruff check --fix`). Review the updated diff, "
|
||||
"then continue or wrap up.",
|
||||
"Repo hook: ran targeted Ruff tidy on touched files.",
|
||||
)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -17,7 +17,7 @@ A clear and concise description of what the bug is.
|
||||
|
||||
### Debug information
|
||||
- Agents SDK version: (e.g. `v0.0.3`)
|
||||
- Python version (e.g. Python 3.10)
|
||||
- Python version (e.g. Python 3.14)
|
||||
|
||||
### Repro steps
|
||||
|
||||
|
||||
@@ -17,7 +17,7 @@ A clear and concise description of what the question or bug is.
|
||||
|
||||
### Debug information
|
||||
- Agents SDK version: (e.g. `v0.0.3`)
|
||||
- Python version (e.g. Python 3.10)
|
||||
- Python version (e.g. Python 3.14)
|
||||
|
||||
### Repro steps
|
||||
Ideally provide a minimal python script that can be run to reproduce the issue.
|
||||
|
||||
@@ -0,0 +1,9 @@
|
||||
version: 2
|
||||
updates:
|
||||
- package-ecosystem: "github-actions"
|
||||
directory: "/"
|
||||
schedule:
|
||||
interval: "monthly"
|
||||
open-pull-requests-limit: 5
|
||||
labels:
|
||||
- "dependencies"
|
||||
Executable
+63
@@ -0,0 +1,63 @@
|
||||
#!/usr/bin/env bash
|
||||
set -euo pipefail
|
||||
|
||||
mode="${1:-code}"
|
||||
base_sha="${2:-${BASE_SHA:-}}"
|
||||
head_sha="${3:-${HEAD_SHA:-}}"
|
||||
|
||||
if [ -z "${GITHUB_OUTPUT:-}" ]; then
|
||||
echo "GITHUB_OUTPUT is not set." >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
if [ -z "$head_sha" ]; then
|
||||
head_sha="$(git rev-parse HEAD 2>/dev/null || true)"
|
||||
fi
|
||||
|
||||
if [ -z "$base_sha" ]; then
|
||||
if ! git rev-parse --verify origin/main >/dev/null 2>&1; then
|
||||
git fetch --no-tags --depth=1 origin main || true
|
||||
fi
|
||||
if git rev-parse --verify origin/main >/dev/null 2>&1 && [ -n "$head_sha" ]; then
|
||||
base_sha="$(git merge-base origin/main "$head_sha" 2>/dev/null || true)"
|
||||
fi
|
||||
fi
|
||||
|
||||
if [ -z "$base_sha" ] || [ -z "$head_sha" ]; then
|
||||
echo "run=true" >> "$GITHUB_OUTPUT"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
if [ "$base_sha" = "0000000000000000000000000000000000000000" ]; then
|
||||
echo "run=true" >> "$GITHUB_OUTPUT"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
if ! git cat-file -e "$base_sha" 2>/dev/null; then
|
||||
git fetch --no-tags --depth=1 origin "$base_sha" || true
|
||||
fi
|
||||
|
||||
if ! git cat-file -e "$base_sha" 2>/dev/null; then
|
||||
echo "run=true" >> "$GITHUB_OUTPUT"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
changed_files=$(git diff --name-only "$base_sha" "$head_sha" || true)
|
||||
|
||||
case "$mode" in
|
||||
code)
|
||||
pattern='^(src/|tests/|examples/|pyproject.toml$|uv.lock$|Makefile$)'
|
||||
;;
|
||||
docs)
|
||||
pattern='^(docs/|mkdocs.yml$)'
|
||||
;;
|
||||
*)
|
||||
pattern="$mode"
|
||||
;;
|
||||
esac
|
||||
|
||||
if echo "$changed_files" | grep -Eq "$pattern"; then
|
||||
echo "run=true" >> "$GITHUB_OUTPUT"
|
||||
else
|
||||
echo "run=false" >> "$GITHUB_OUTPUT"
|
||||
fi
|
||||
@@ -0,0 +1,20 @@
|
||||
#!/usr/bin/env bash
|
||||
set -euo pipefail
|
||||
|
||||
repeat_count="${1:-5}"
|
||||
|
||||
asyncio_progress_args=(
|
||||
tests/test_asyncio_progress.py
|
||||
)
|
||||
|
||||
run_step_execution_args=(
|
||||
tests/test_run_step_execution.py
|
||||
-k
|
||||
"cancel or post_invoke"
|
||||
)
|
||||
|
||||
for run in $(seq 1 "$repeat_count"); do
|
||||
echo "Async teardown stability run ${run}/${repeat_count}"
|
||||
uv run pytest -q "${asyncio_progress_args[@]}"
|
||||
uv run pytest -q "${run_step_execution_args[@]}"
|
||||
done
|
||||
@@ -0,0 +1,184 @@
|
||||
#!/usr/bin/env python3
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import os
|
||||
import re
|
||||
import subprocess
|
||||
import sys
|
||||
from urllib import error, request
|
||||
|
||||
|
||||
def warn(message: str) -> None:
|
||||
print(message, file=sys.stderr)
|
||||
|
||||
|
||||
def parse_version(value: str | None) -> tuple[int, int, int] | None:
|
||||
if not value:
|
||||
return None
|
||||
match = re.match(r"^v?(\d+)\.(\d+)(?:\.(\d+))?", value)
|
||||
if not match:
|
||||
return None
|
||||
major = int(match.group(1))
|
||||
minor = int(match.group(2))
|
||||
patch = int(match.group(3) or 0)
|
||||
return major, minor, patch
|
||||
|
||||
|
||||
def latest_tag_version(exclude_version: tuple[int, int, int] | None) -> tuple[int, int, int] | None:
|
||||
try:
|
||||
output = subprocess.check_output(["git", "tag", "--list", "v*"], text=True)
|
||||
except Exception as exc:
|
||||
warn(f"Milestone assignment skipped (failed to list tags: {exc}).")
|
||||
return None
|
||||
versions: list[tuple[int, int, int]] = []
|
||||
for tag in output.splitlines():
|
||||
parsed = parse_version(tag)
|
||||
if not parsed:
|
||||
continue
|
||||
if exclude_version and parsed == exclude_version:
|
||||
continue
|
||||
versions.append(parsed)
|
||||
if not versions:
|
||||
return None
|
||||
return max(versions)
|
||||
|
||||
|
||||
def classify_bump(
|
||||
target: tuple[int, int, int] | None,
|
||||
previous: tuple[int, int, int] | None,
|
||||
) -> str | None:
|
||||
if not target or not previous:
|
||||
return None
|
||||
if target < previous:
|
||||
warn("Milestone assignment skipped (release version is behind latest tag).")
|
||||
return None
|
||||
if target[0] != previous[0]:
|
||||
return "major"
|
||||
if target[1] != previous[1]:
|
||||
return "minor"
|
||||
return "patch"
|
||||
|
||||
|
||||
def parse_milestone_title(title: str | None) -> tuple[int, int] | None:
|
||||
if not title:
|
||||
return None
|
||||
match = re.match(r"^(\d+)\.(\d+)\.x$", title)
|
||||
if not match:
|
||||
return None
|
||||
return int(match.group(1)), int(match.group(2))
|
||||
|
||||
|
||||
def fetch_open_milestones(owner: str, repo: str, token: str) -> list[dict]:
|
||||
url = f"https://api.github.com/repos/{owner}/{repo}/milestones?state=open&per_page=100"
|
||||
headers = {
|
||||
"Accept": "application/vnd.github+json",
|
||||
"Authorization": f"Bearer {token}",
|
||||
}
|
||||
req = request.Request(url, headers=headers)
|
||||
try:
|
||||
with request.urlopen(req) as response:
|
||||
return json.load(response)
|
||||
except error.HTTPError as exc:
|
||||
warn(f"Milestone assignment skipped (failed to list milestones: {exc.code}).")
|
||||
except Exception as exc:
|
||||
warn(f"Milestone assignment skipped (failed to list milestones: {exc}).")
|
||||
return []
|
||||
|
||||
|
||||
def select_milestone(milestones: list[dict], required_bump: str) -> str | None:
|
||||
parsed: list[dict] = []
|
||||
for milestone in milestones:
|
||||
parsed_title = parse_milestone_title(milestone.get("title"))
|
||||
if not parsed_title:
|
||||
continue
|
||||
parsed.append(
|
||||
{
|
||||
"milestone": milestone,
|
||||
"major": parsed_title[0],
|
||||
"minor": parsed_title[1],
|
||||
}
|
||||
)
|
||||
|
||||
parsed.sort(key=lambda entry: (entry["major"], entry["minor"]))
|
||||
if not parsed:
|
||||
warn("Milestone assignment skipped (no open milestones matching X.Y.x).")
|
||||
return None
|
||||
|
||||
majors = sorted({entry["major"] for entry in parsed})
|
||||
current_major = majors[0]
|
||||
next_major = majors[1] if len(majors) > 1 else None
|
||||
|
||||
current_major_entries = [entry for entry in parsed if entry["major"] == current_major]
|
||||
patch_target = current_major_entries[0]
|
||||
minor_target = current_major_entries[1] if len(current_major_entries) > 1 else patch_target
|
||||
|
||||
major_target = None
|
||||
if next_major is not None:
|
||||
next_major_entries = [entry for entry in parsed if entry["major"] == next_major]
|
||||
if next_major_entries:
|
||||
major_target = next_major_entries[0]
|
||||
|
||||
target_entry = None
|
||||
if required_bump == "major":
|
||||
target_entry = major_target
|
||||
elif required_bump == "minor":
|
||||
target_entry = minor_target
|
||||
else:
|
||||
target_entry = patch_target
|
||||
|
||||
if not target_entry:
|
||||
warn("Milestone assignment skipped (not enough open milestones for selection).")
|
||||
return None
|
||||
|
||||
return target_entry["milestone"].get("title")
|
||||
|
||||
|
||||
def main() -> int:
|
||||
parser = argparse.ArgumentParser()
|
||||
parser.add_argument("--version", help="Release version (e.g., 0.6.6).")
|
||||
parser.add_argument(
|
||||
"--required-bump",
|
||||
choices=("major", "minor", "patch"),
|
||||
help="Override bump type (major/minor/patch).",
|
||||
)
|
||||
parser.add_argument("--repo", help="GitHub repository (owner/repo).")
|
||||
parser.add_argument("--token", help="GitHub token.")
|
||||
args = parser.parse_args()
|
||||
|
||||
required_bump = args.required_bump
|
||||
if not required_bump:
|
||||
target_version = parse_version(args.version)
|
||||
if not target_version:
|
||||
warn("Milestone assignment skipped (missing or invalid release version).")
|
||||
return 0
|
||||
previous_version = latest_tag_version(target_version)
|
||||
required_bump = classify_bump(target_version, previous_version)
|
||||
if not required_bump:
|
||||
warn("Milestone assignment skipped (unable to determine required bump).")
|
||||
return 0
|
||||
|
||||
token = args.token or os.environ.get("GITHUB_TOKEN") or os.environ.get("GH_TOKEN")
|
||||
if not token:
|
||||
warn("Milestone assignment skipped (missing GitHub token).")
|
||||
return 0
|
||||
|
||||
repo = args.repo or os.environ.get("GITHUB_REPOSITORY")
|
||||
if not repo or "/" not in repo:
|
||||
warn("Milestone assignment skipped (missing repository info).")
|
||||
return 0
|
||||
owner, name = repo.split("/", 1)
|
||||
|
||||
milestones = fetch_open_milestones(owner, name, token)
|
||||
if not milestones:
|
||||
return 0
|
||||
|
||||
milestone_title = select_milestone(milestones, required_bump)
|
||||
if milestone_title:
|
||||
print(milestone_title)
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
@@ -1,26 +1,48 @@
|
||||
name: Deploy docs
|
||||
|
||||
on:
|
||||
workflow_run:
|
||||
workflows: ["Tests"]
|
||||
types:
|
||||
- completed
|
||||
push:
|
||||
branches:
|
||||
- main
|
||||
paths:
|
||||
- "docs/**"
|
||||
- "mkdocs.yml"
|
||||
|
||||
permissions:
|
||||
contents: write # This allows pushing to gh-pages
|
||||
|
||||
jobs:
|
||||
deploy_docs:
|
||||
if: ${{ github.event.workflow_run.conclusion == 'success' }}
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
uses: actions/checkout@v4
|
||||
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd
|
||||
- name: Determine docs-only push
|
||||
id: docs-only
|
||||
run: |
|
||||
if [ "${{ github.event_name }}" != "push" ]; then
|
||||
echo "skip=false" >> "$GITHUB_OUTPUT"
|
||||
exit 0
|
||||
fi
|
||||
set -euo pipefail
|
||||
before="${{ github.event.before }}"
|
||||
sha="${{ github.sha }}"
|
||||
changed_files=$(git diff --name-only "$before" "$sha" || true)
|
||||
non_docs=$(echo "$changed_files" | grep -vE '^(docs/|mkdocs.yml$)' || true)
|
||||
if [ -n "$non_docs" ]; then
|
||||
echo "skip=true" >> "$GITHUB_OUTPUT"
|
||||
else
|
||||
echo "skip=false" >> "$GITHUB_OUTPUT"
|
||||
fi
|
||||
- name: Setup uv
|
||||
uses: astral-sh/setup-uv@v5
|
||||
if: steps.docs-only.outputs.skip != 'true'
|
||||
uses: astral-sh/setup-uv@08807647e7069bb48b6ef5acd8ec9567f424441b # setup-uv v8.1.0; uv 0.11.14
|
||||
with:
|
||||
version: "0.11.14"
|
||||
enable-cache: true
|
||||
- name: Install dependencies
|
||||
if: steps.docs-only.outputs.skip != 'true'
|
||||
run: make sync
|
||||
- name: Deploy docs
|
||||
if: steps.docs-only.outputs.skip != 'true'
|
||||
run: make deploy-docs
|
||||
|
||||
@@ -10,17 +10,19 @@ jobs:
|
||||
issues: write
|
||||
pull-requests: write
|
||||
steps:
|
||||
- uses: actions/stale@v9
|
||||
- uses: actions/stale@b5d41d4e1d5dceea10e7104786b73624c18a190f
|
||||
with:
|
||||
days-before-issue-stale: 7
|
||||
days-before-issue-close: 3
|
||||
stale-issue-label: "stale"
|
||||
exempt-issue-labels: "skip-stale"
|
||||
stale-issue-message: "This issue is stale because it has been open for 7 days with no activity."
|
||||
close-issue-message: "This issue was closed because it has been inactive for 3 days since being marked as stale."
|
||||
any-of-issue-labels: 'question,needs-more-info'
|
||||
days-before-pr-stale: 10
|
||||
days-before-pr-close: 7
|
||||
stale-pr-label: "stale"
|
||||
exempt-pr-labels: "skip-stale"
|
||||
stale-pr-message: "This PR is stale because it has been open for 10 days with no activity."
|
||||
close-pr-message: "This PR was closed because it has been inactive for 7 days since being marked as stale."
|
||||
repo-token: ${{ secrets.GITHUB_TOKEN }}
|
||||
|
||||
@@ -21,14 +21,15 @@ jobs:
|
||||
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
uses: actions/checkout@v4
|
||||
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd
|
||||
- name: Setup uv
|
||||
uses: astral-sh/setup-uv@v5
|
||||
uses: astral-sh/setup-uv@08807647e7069bb48b6ef5acd8ec9567f424441b # setup-uv v8.1.0; uv 0.11.14
|
||||
with:
|
||||
version: "0.11.14"
|
||||
enable-cache: true
|
||||
- name: Install dependencies
|
||||
run: make sync
|
||||
- name: Build package
|
||||
run: uv build
|
||||
- name: Publish to PyPI
|
||||
uses: pypa/gh-action-pypi-publish@release/v1
|
||||
uses: pypa/gh-action-pypi-publish@cef221092ed1bacb1cc03d23a2d87d1d172e277b
|
||||
|
||||
@@ -0,0 +1,143 @@
|
||||
name: Create release PR
|
||||
|
||||
on:
|
||||
workflow_dispatch:
|
||||
inputs:
|
||||
version:
|
||||
description: "Version to release (e.g., 0.6.6)"
|
||||
required: true
|
||||
|
||||
permissions:
|
||||
contents: write
|
||||
pull-requests: write
|
||||
|
||||
jobs:
|
||||
release-pr:
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd
|
||||
with:
|
||||
fetch-depth: 0
|
||||
ref: main
|
||||
- name: Setup uv
|
||||
uses: astral-sh/setup-uv@08807647e7069bb48b6ef5acd8ec9567f424441b # setup-uv v8.1.0; uv 0.11.14
|
||||
with:
|
||||
version: "0.11.14"
|
||||
enable-cache: true
|
||||
- name: Fetch tags
|
||||
run: git fetch origin --tags --prune
|
||||
- name: Ensure release branch does not exist
|
||||
env:
|
||||
RELEASE_VERSION: ${{ inputs.version }}
|
||||
run: |
|
||||
branch="release/v${RELEASE_VERSION}"
|
||||
if git ls-remote --exit-code --heads origin "$branch" >/dev/null 2>&1; then
|
||||
echo "Branch $branch already exists on origin." >&2
|
||||
exit 1
|
||||
fi
|
||||
- name: Update version
|
||||
env:
|
||||
RELEASE_VERSION: ${{ inputs.version }}
|
||||
run: |
|
||||
python - <<'PY'
|
||||
import os
|
||||
import pathlib
|
||||
import re
|
||||
import sys
|
||||
|
||||
version = os.environ["RELEASE_VERSION"]
|
||||
if version.startswith("v"):
|
||||
print("Version must not start with 'v' (use x.y.z...).", file=sys.stderr)
|
||||
sys.exit(1)
|
||||
if ".." in version:
|
||||
print("Version contains consecutive dots (use x.y.z...).", file=sys.stderr)
|
||||
sys.exit(1)
|
||||
if not re.match(r"^\d+\.\d+(\.\d+)*([a-zA-Z0-9\.-]+)?$", version):
|
||||
print(
|
||||
"Version must be semver-like (e.g., 0.6.6, 0.6.6-rc1, 0.6.6.dev1).",
|
||||
file=sys.stderr,
|
||||
)
|
||||
sys.exit(1)
|
||||
path = pathlib.Path("pyproject.toml")
|
||||
text = path.read_text()
|
||||
updated, count = re.subn(
|
||||
r'(?m)^version\s*=\s*"[^\"]+"',
|
||||
f'version = "{version}"',
|
||||
text,
|
||||
)
|
||||
if count != 1:
|
||||
print("Expected to update exactly one version line.", file=sys.stderr)
|
||||
sys.exit(1)
|
||||
if updated == text:
|
||||
print("Version already set; no changes made.", file=sys.stderr)
|
||||
sys.exit(1)
|
||||
path.write_text(updated)
|
||||
PY
|
||||
- name: Sync dependencies
|
||||
run: make sync
|
||||
- name: Configure git
|
||||
run: |
|
||||
git config user.name "github-actions[bot]"
|
||||
git config user.email "github-actions[bot]@users.noreply.github.com"
|
||||
- name: Create release branch and commit
|
||||
env:
|
||||
RELEASE_VERSION: ${{ inputs.version }}
|
||||
run: |
|
||||
branch="release/v${RELEASE_VERSION}"
|
||||
git checkout -b "$branch"
|
||||
git add pyproject.toml uv.lock
|
||||
if git diff --cached --quiet; then
|
||||
echo "No changes to commit." >&2
|
||||
exit 1
|
||||
fi
|
||||
git commit -m "Bump version to ${RELEASE_VERSION}"
|
||||
git push --set-upstream origin "$branch"
|
||||
- name: Build PR body
|
||||
env:
|
||||
RELEASE_VERSION: ${{ inputs.version }}
|
||||
run: |
|
||||
printf 'Release PR for %s.\n\nThe release readiness report will be prepared manually.\n' "$RELEASE_VERSION" > pr-body.md
|
||||
- name: Create or update PR
|
||||
env:
|
||||
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
RELEASE_VERSION: ${{ inputs.version }}
|
||||
run: |
|
||||
set -euo pipefail
|
||||
head_branch="release/v${RELEASE_VERSION}"
|
||||
milestone_name="$(python .github/scripts/select-release-milestone.py --version "$RELEASE_VERSION")"
|
||||
pr_number="$(gh pr list --head "$head_branch" --base "main" --json number --jq '.[0].number // empty')"
|
||||
if [ -z "$pr_number" ]; then
|
||||
create_args=(
|
||||
--title "Release ${RELEASE_VERSION}"
|
||||
--body-file pr-body.md
|
||||
--base "main"
|
||||
--head "$head_branch"
|
||||
--label "project"
|
||||
)
|
||||
if [ -n "$milestone_name" ]; then
|
||||
create_args+=(--milestone "$milestone_name")
|
||||
fi
|
||||
if ! gh pr create "${create_args[@]}"; then
|
||||
echo "PR create with label/milestone failed; retrying without them." >&2
|
||||
gh pr create \
|
||||
--title "Release ${RELEASE_VERSION}" \
|
||||
--body-file pr-body.md \
|
||||
--base "main" \
|
||||
--head "$head_branch"
|
||||
fi
|
||||
else
|
||||
edit_args=(
|
||||
--title "Release ${RELEASE_VERSION}"
|
||||
--body-file pr-body.md
|
||||
--add-label "project"
|
||||
)
|
||||
if [ -n "$milestone_name" ]; then
|
||||
edit_args+=(--milestone "$milestone_name")
|
||||
fi
|
||||
if ! gh pr edit "$pr_number" "${edit_args[@]}"; then
|
||||
echo "PR edit with label/milestone failed; retrying without them." >&2
|
||||
gh pr edit "$pr_number" --title "Release ${RELEASE_VERSION}" --body-file pr-body.md
|
||||
fi
|
||||
fi
|
||||
@@ -0,0 +1,84 @@
|
||||
name: Tag release on merge
|
||||
|
||||
on:
|
||||
pull_request:
|
||||
types:
|
||||
- closed
|
||||
branches:
|
||||
- main
|
||||
|
||||
permissions:
|
||||
contents: write
|
||||
|
||||
jobs:
|
||||
tag-release:
|
||||
if: >-
|
||||
github.event.pull_request.merged == true &&
|
||||
github.event.pull_request.head.repo.full_name == github.repository &&
|
||||
startsWith(github.event.pull_request.head.ref, 'release/v')
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- name: Validate merge commit
|
||||
env:
|
||||
MERGE_SHA: ${{ github.event.pull_request.merge_commit_sha }}
|
||||
run: |
|
||||
if [ -z "$MERGE_SHA" ]; then
|
||||
echo "merge_commit_sha is empty; refusing to tag to avoid tagging the wrong commit." >&2
|
||||
exit 1
|
||||
fi
|
||||
- name: Checkout merge commit
|
||||
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd
|
||||
with:
|
||||
fetch-depth: 0
|
||||
ref: ${{ github.event.pull_request.merge_commit_sha }}
|
||||
- name: Setup Python
|
||||
uses: actions/setup-python@a309ff8b426b58ec0e2a45f0f869d46889d02405
|
||||
with:
|
||||
python-version: "3.11"
|
||||
- name: Configure git
|
||||
run: |
|
||||
git config user.name "github-actions[bot]"
|
||||
git config user.email "github-actions[bot]@users.noreply.github.com"
|
||||
- name: Fetch tags
|
||||
run: git fetch origin --tags --prune
|
||||
- name: Resolve version
|
||||
id: version
|
||||
env:
|
||||
HEAD_REF: ${{ github.event.pull_request.head.ref }}
|
||||
run: |
|
||||
python - <<'PY'
|
||||
import os
|
||||
import pathlib
|
||||
import sys
|
||||
import tomllib
|
||||
|
||||
path = pathlib.Path("pyproject.toml")
|
||||
data = tomllib.loads(path.read_text())
|
||||
version = data.get("project", {}).get("version")
|
||||
if not version:
|
||||
print("Missing project.version in pyproject.toml.", file=sys.stderr)
|
||||
sys.exit(1)
|
||||
|
||||
head_ref = os.environ.get("HEAD_REF", "")
|
||||
if head_ref.startswith("release/v"):
|
||||
expected = head_ref[len("release/v") :]
|
||||
if expected != version:
|
||||
print(
|
||||
f"Version mismatch: branch {expected} vs pyproject {version}.",
|
||||
file=sys.stderr,
|
||||
)
|
||||
sys.exit(1)
|
||||
|
||||
output_path = pathlib.Path(os.environ["GITHUB_OUTPUT"])
|
||||
output_path.write_text(f"version={version}\n")
|
||||
PY
|
||||
- name: Create tag
|
||||
env:
|
||||
VERSION: ${{ steps.version.outputs.version }}
|
||||
run: |
|
||||
if git tag -l "v${VERSION}" | grep -q "v${VERSION}"; then
|
||||
echo "Tag v${VERSION} already exists; skipping."
|
||||
exit 0
|
||||
fi
|
||||
git tag -a "v${VERSION}" -m "Release v${VERSION}"
|
||||
git push origin "v${VERSION}"
|
||||
+100
-27
@@ -5,8 +5,10 @@ on:
|
||||
branches:
|
||||
- main
|
||||
pull_request:
|
||||
branches:
|
||||
- main
|
||||
# All PRs, including stacked PRs
|
||||
|
||||
permissions:
|
||||
contents: read
|
||||
|
||||
env:
|
||||
UV_FROZEN: "1"
|
||||
@@ -16,45 +18,122 @@ jobs:
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
uses: actions/checkout@v4
|
||||
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd
|
||||
- name: Detect code changes
|
||||
id: changes
|
||||
run: ./.github/scripts/detect-changes.sh code "${{ github.event.pull_request.base.sha || github.event.before }}" "${{ github.sha }}"
|
||||
- name: Setup uv
|
||||
uses: astral-sh/setup-uv@v5
|
||||
if: steps.changes.outputs.run == 'true'
|
||||
uses: astral-sh/setup-uv@08807647e7069bb48b6ef5acd8ec9567f424441b # setup-uv v8.1.0; uv 0.11.14
|
||||
with:
|
||||
version: "0.11.14"
|
||||
enable-cache: true
|
||||
- name: Install dependencies
|
||||
if: steps.changes.outputs.run == 'true'
|
||||
run: make sync
|
||||
- name: Verify formatting
|
||||
if: steps.changes.outputs.run == 'true'
|
||||
run: make format-check
|
||||
- name: Run lint
|
||||
if: steps.changes.outputs.run == 'true'
|
||||
run: make lint
|
||||
- name: Skip lint
|
||||
if: steps.changes.outputs.run != 'true'
|
||||
run: echo "Skipping lint for non-code changes."
|
||||
|
||||
typecheck:
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
uses: actions/checkout@v4
|
||||
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd
|
||||
- name: Detect code changes
|
||||
id: changes
|
||||
run: ./.github/scripts/detect-changes.sh code "${{ github.event.pull_request.base.sha || github.event.before }}" "${{ github.sha }}"
|
||||
- name: Setup uv
|
||||
uses: astral-sh/setup-uv@v5
|
||||
if: steps.changes.outputs.run == 'true'
|
||||
uses: astral-sh/setup-uv@08807647e7069bb48b6ef5acd8ec9567f424441b # setup-uv v8.1.0; uv 0.11.14
|
||||
with:
|
||||
version: "0.11.14"
|
||||
enable-cache: true
|
||||
- name: Install dependencies
|
||||
if: steps.changes.outputs.run == 'true'
|
||||
run: make sync
|
||||
- name: Run typecheck
|
||||
run: make mypy
|
||||
if: steps.changes.outputs.run == 'true'
|
||||
run: make typecheck
|
||||
- name: Skip typecheck
|
||||
if: steps.changes.outputs.run != 'true'
|
||||
run: echo "Skipping typecheck for non-code changes."
|
||||
|
||||
tests:
|
||||
runs-on: ubuntu-latest
|
||||
strategy:
|
||||
fail-fast: false
|
||||
matrix:
|
||||
python-version:
|
||||
- "3.10"
|
||||
- "3.11"
|
||||
- "3.12"
|
||||
- "3.13"
|
||||
- "3.14"
|
||||
env:
|
||||
OPENAI_API_KEY: fake-for-tests
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
uses: actions/checkout@v4
|
||||
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd
|
||||
- name: Detect code changes
|
||||
id: changes
|
||||
run: ./.github/scripts/detect-changes.sh code "${{ github.event.pull_request.base.sha || github.event.before }}" "${{ github.sha }}"
|
||||
- name: Setup uv
|
||||
uses: astral-sh/setup-uv@v5
|
||||
if: steps.changes.outputs.run == 'true'
|
||||
uses: astral-sh/setup-uv@08807647e7069bb48b6ef5acd8ec9567f424441b # setup-uv v8.1.0; uv 0.11.14
|
||||
with:
|
||||
version: "0.11.14"
|
||||
enable-cache: true
|
||||
python-version: ${{ matrix.python-version }}
|
||||
- name: Install dependencies
|
||||
if: steps.changes.outputs.run == 'true'
|
||||
run: make sync
|
||||
- name: Run tests with coverage
|
||||
if: steps.changes.outputs.run == 'true' && matrix.python-version == '3.12'
|
||||
run: make coverage
|
||||
- name: Run tests
|
||||
if: steps.changes.outputs.run == 'true' && matrix.python-version != '3.12'
|
||||
run: make tests
|
||||
- name: Run async teardown stability tests
|
||||
if: steps.changes.outputs.run == 'true' && (matrix.python-version == '3.10' || matrix.python-version == '3.14')
|
||||
run: make tests-asyncio-stability
|
||||
- name: Skip tests
|
||||
if: steps.changes.outputs.run != 'true'
|
||||
run: echo "Skipping tests for non-code changes."
|
||||
|
||||
tests-windows:
|
||||
runs-on: windows-latest
|
||||
env:
|
||||
OPENAI_API_KEY: fake-for-tests
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd
|
||||
- name: Detect code changes
|
||||
id: changes
|
||||
shell: bash
|
||||
run: ./.github/scripts/detect-changes.sh code "${{ github.event.pull_request.base.sha || github.event.before }}" "${{ github.sha }}"
|
||||
- name: Setup uv
|
||||
if: steps.changes.outputs.run == 'true'
|
||||
uses: astral-sh/setup-uv@08807647e7069bb48b6ef5acd8ec9567f424441b # setup-uv v8.1.0; uv 0.11.14
|
||||
with:
|
||||
version: "0.11.14"
|
||||
enable-cache: true
|
||||
python-version: "3.13"
|
||||
- name: Install dependencies
|
||||
if: steps.changes.outputs.run == 'true'
|
||||
run: uv sync --all-extras --all-packages --group dev
|
||||
- name: Run tests
|
||||
if: steps.changes.outputs.run == 'true'
|
||||
run: uv run pytest
|
||||
- name: Skip tests
|
||||
if: steps.changes.outputs.run != 'true'
|
||||
run: echo "Skipping tests for non-code changes."
|
||||
|
||||
build-docs:
|
||||
runs-on: ubuntu-latest
|
||||
@@ -62,28 +141,22 @@ jobs:
|
||||
OPENAI_API_KEY: fake-for-tests
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
uses: actions/checkout@v4
|
||||
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd
|
||||
- name: Detect docs changes
|
||||
id: changes
|
||||
run: ./.github/scripts/detect-changes.sh docs "${{ github.event.pull_request.base.sha || github.event.before }}" "${{ github.sha }}"
|
||||
- name: Setup uv
|
||||
uses: astral-sh/setup-uv@v5
|
||||
if: steps.changes.outputs.run == 'true'
|
||||
uses: astral-sh/setup-uv@08807647e7069bb48b6ef5acd8ec9567f424441b # setup-uv v8.1.0; uv 0.11.14
|
||||
with:
|
||||
version: "0.11.14"
|
||||
enable-cache: true
|
||||
- name: Install dependencies
|
||||
if: steps.changes.outputs.run == 'true'
|
||||
run: make sync
|
||||
- name: Build docs
|
||||
if: steps.changes.outputs.run == 'true'
|
||||
run: make build-docs
|
||||
|
||||
old_versions:
|
||||
runs-on: ubuntu-latest
|
||||
env:
|
||||
OPENAI_API_KEY: fake-for-tests
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
uses: actions/checkout@v4
|
||||
- name: Setup uv
|
||||
uses: astral-sh/setup-uv@v5
|
||||
with:
|
||||
enable-cache: true
|
||||
- name: Install dependencies
|
||||
run: make sync
|
||||
- name: Run tests
|
||||
run: make old_version_tests
|
||||
- name: Skip docs build
|
||||
if: steps.changes.outputs.run != 'true'
|
||||
run: echo "Skipping docs build for non-docs changes."
|
||||
|
||||
+19
-2
@@ -45,6 +45,7 @@ htmlcov/
|
||||
.coverage
|
||||
.coverage.*
|
||||
.cache
|
||||
.tmp/
|
||||
nosetests.xml
|
||||
coverage.xml
|
||||
*.cover
|
||||
@@ -100,8 +101,10 @@ celerybeat.pid
|
||||
*.sage.py
|
||||
|
||||
# Environments
|
||||
.env
|
||||
.python-version
|
||||
.env*
|
||||
.venv
|
||||
.venv*
|
||||
env/
|
||||
venv/
|
||||
ENV/
|
||||
@@ -140,5 +143,19 @@ cython_debug/
|
||||
# Ruff stuff:
|
||||
.ruff_cache/
|
||||
|
||||
# Example runtime state
|
||||
examples/sandbox/extensions/daytona/usaspending_text2sql/.audit_log.jsonl
|
||||
examples/sandbox/extensions/daytona/usaspending_text2sql/.session_state.json
|
||||
|
||||
# PyPI configuration file
|
||||
.pypirc
|
||||
.pypirc
|
||||
.aider*
|
||||
|
||||
# Redis database files
|
||||
dump.rdb
|
||||
|
||||
tmp/
|
||||
|
||||
# execplans
|
||||
plans/
|
||||
.vercel
|
||||
|
||||
Vendored
+14
@@ -0,0 +1,14 @@
|
||||
{
|
||||
// Use IntelliSense to learn about possible attributes.
|
||||
// Hover to view descriptions of existing attributes.
|
||||
// For more information, visit: https://go.microsoft.com/fwlink/?linkid=830387
|
||||
"version": "0.2.0",
|
||||
"configurations": [
|
||||
{
|
||||
"name": "Python Debugger: Python File",
|
||||
"type": "debugpy",
|
||||
"request": "launch",
|
||||
"program": "${file}"
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,231 @@
|
||||
# Contributor Guide
|
||||
|
||||
This guide helps new contributors get started with the OpenAI Agents Python repository. It covers repo structure, how to test your work, available utilities, and guidelines for commits and PRs.
|
||||
|
||||
**Location:** `AGENTS.md` at the repository root.
|
||||
|
||||
## Table of Contents
|
||||
|
||||
1. [Policies & Mandatory Rules](#policies--mandatory-rules)
|
||||
2. [Project Structure Guide](#project-structure-guide)
|
||||
3. [Operation Guide](#operation-guide)
|
||||
|
||||
## Policies & Mandatory Rules
|
||||
|
||||
### Mandatory Skill Usage
|
||||
|
||||
#### `$code-change-verification`
|
||||
|
||||
Run `$code-change-verification` before marking work complete when changes affect runtime code, tests, or build/test behavior.
|
||||
|
||||
Run it when you change:
|
||||
- `src/agents/` (library code) or shared utilities.
|
||||
- `tests/` or add or modify snapshot tests.
|
||||
- `examples/`.
|
||||
- Build or test configuration such as `pyproject.toml`, `Makefile`, `mkdocs.yml`, `docs/scripts/`, or CI workflows.
|
||||
|
||||
You can skip `$code-change-verification` for docs-only or repo-meta changes (for example, `docs/`, `.agents/`, `README.md`, `AGENTS.md`, `.github/`), unless a user explicitly asks to run the full verification stack.
|
||||
|
||||
#### `$openai-knowledge`
|
||||
|
||||
When working on OpenAI API or OpenAI platform integrations in this repo (Responses API, tools, streaming, Realtime API, auth, models, rate limits, MCP, Agents SDK or ChatGPT Apps SDK), use `$openai-knowledge` to pull authoritative docs via the OpenAI Developer Docs MCP server (and guide setup if it is not configured).
|
||||
|
||||
#### `$implementation-strategy`
|
||||
|
||||
Before changing runtime code, exported APIs, external configuration, persisted schemas, wire protocols, or other user-facing behavior, use `$implementation-strategy` to decide the compatibility boundary and implementation shape. Judge breaking changes against the latest release tag, not unreleased branch-local churn. Interfaces introduced or changed after the latest release tag may be rewritten without compatibility shims unless they define a released or explicitly supported durable external state boundary, or the user explicitly asks for a migration path. Unreleased persisted formats on `main` may be renumbered or squashed before release when intermediate snapshots are intentionally unsupported.
|
||||
|
||||
#### `$pr-draft-summary`
|
||||
|
||||
When a task in this repo finishes with moderate-or-larger code changes, invoke `$pr-draft-summary` in the final handoff to generate the required PR summary block, branch suggestion, title, and draft description. Treat this as the default close-out step after runtime code, tests, examples, build/test configuration, or docs with behavior impact are changed.
|
||||
|
||||
Skip `$pr-draft-summary` only for trivial or conversation-only tasks, repo-meta/doc-only tasks without behavior impact, or when the user explicitly says not to include the PR draft block.
|
||||
|
||||
### ExecPlans
|
||||
|
||||
Call out compatibility risk early in your plan only when the change affects behavior shipped in the latest release tag or a released or explicitly supported durable external state boundary, and confirm the approach before implementing changes that could impact users.
|
||||
|
||||
Use an ExecPlan when work is multi-step, spans several files, involves new features or refactors, or is likely to take more than about an hour. Start with the template and rules in `PLANS.md`, keep milestones and living sections (Progress, Surprises & Discoveries, Decision Log, Outcomes & Retrospective) up to date as you execute, and rewrite the plan if scope shifts. Call out compatibility risk only when the plan changes behavior shipped in the latest release tag or a released or explicitly supported durable external state boundary. Do not treat branch-local interface churn or unreleased post-tag changes on `main` as breaking by default; prefer direct replacement over compatibility layers in those cases, and renumber or squash unreleased persisted schemas before release when the intermediate snapshots are intentionally unsupported. If you intentionally skip an ExecPlan for a complex task, note why in your response so reviewers understand the choice.
|
||||
|
||||
### Public API Positional Compatibility
|
||||
|
||||
Treat the parameter and dataclass field order of exported runtime APIs as a compatibility contract.
|
||||
|
||||
- For public constructors (for example `RunConfig`, `FunctionTool`, `AgentHookContext`), preserve existing positional argument meaning. Do not insert new constructor parameters or dataclass fields in the middle of existing public order.
|
||||
- When adding a new optional public field/parameter, append it to the end whenever possible and keep old fields in the same order.
|
||||
- If reordering is unavoidable, add an explicit compatibility layer and regression tests that exercise the old positional call pattern.
|
||||
- Prefer keyword arguments at call sites to reduce accidental breakage, but do not rely on this to justify breaking positional compatibility for public APIs.
|
||||
|
||||
### Platform, Docs, and Security Review
|
||||
|
||||
- Documentation is published to the live site, so coordinate SDK behavior changes and docs carefully. If docs describe behavior that is not released yet, either delay the docs change until the SDK release is available or split it into a follow-up PR.
|
||||
- Treat runnable docs snippets as API compatibility checks. Before adding OpenAI API, provider, Responses, Realtime, WebSocket, or SDK constructor examples, verify the shown arguments and call shape against the actual implementation.
|
||||
- Do not let untrusted sandbox manifests opt themselves out of host filesystem or base-directory boundaries. Escape hatches for local source materialization must be controlled by trusted application code at the call site, not by serialized manifest data.
|
||||
- When documenting sandbox or security grants, verify the actual implementation path enforces the grant or boundary. Do not claim a grant applies to `LocalDir`, `LocalFile`, archive extraction, or other materialization paths unless those paths actually consult it.
|
||||
- When redacting OpenAI tool, MCP, model, or provider payloads, consider traceback display, exception chaining, `__context__`, logs, and telemetry. Suppressing display with `raise ... from None` is not enough if the original exception object still carries sensitive input data.
|
||||
- For OpenAI platform or SDK-specific docs changes, prefer `$openai-knowledge` for authoritative platform behavior and inspect the local code path for SDK behavior. Do not rely on generic API assumptions when documenting Responses, Chat Completions, Realtime, tools, MCP, or provider adapters.
|
||||
|
||||
## Project Structure Guide
|
||||
|
||||
### Overview
|
||||
|
||||
The OpenAI Agents Python repository provides the Python Agents SDK, examples, and documentation built with MkDocs. Use `uv run python ...` for Python commands to ensure a consistent environment.
|
||||
|
||||
### Repo Structure & Important Files
|
||||
|
||||
- `src/agents/`: Core library implementation.
|
||||
- `tests/`: Test suite; see `tests/README.md` for snapshot guidance.
|
||||
- `examples/`: Sample projects showing SDK usage.
|
||||
- `docs/`: MkDocs documentation source; do not edit translated docs under `docs/ja`, `docs/ko`, or `docs/zh` (they are generated).
|
||||
- `docs/scripts/`: Documentation utilities, including translation and reference generation.
|
||||
- `mkdocs.yml`: Documentation site configuration.
|
||||
- `Makefile`: Common developer commands.
|
||||
- `pyproject.toml`, `uv.lock`: Python dependencies and tool configuration.
|
||||
- `.github/PULL_REQUEST_TEMPLATE/pull_request_template.md`: Pull request template to use when opening PRs.
|
||||
- `site/`: Built documentation output.
|
||||
|
||||
### Agents Core Runtime Guidelines
|
||||
|
||||
- `src/agents/run.py` is the runtime entrypoint (`Runner`, `AgentRunner`). Keep it focused on orchestration and public flow control. Put new runtime logic under `src/agents/run_internal/` and import it into `run.py`.
|
||||
- When `run.py` grows, refactor helpers into `run_internal/` modules (for example `run_loop.py`, `turn_resolution.py`, `tool_execution.py`, `session_persistence.py`) and leave only wiring and composition in `run.py`.
|
||||
- Keep streaming and non-streaming paths behaviorally aligned. Changes to `run_internal/run_loop.py` (`run_single_turn`, `run_single_turn_streamed`, `get_new_response`, `start_streaming`) should be mirrored, and any new streaming item types must be reflected in `src/agents/stream_events.py`.
|
||||
- Input guardrails run only on the first turn and only for the starting agent. Resuming an interruption from `RunState` must not increment the turn counter; only actual model calls advance turns.
|
||||
- Server-managed conversation (`conversation_id`, `previous_response_id`, `auto_previous_response_id`) uses `OpenAIServerConversationTracker` in `run_internal/oai_conversation.py`. Only deltas should be sent. If `call_model_input_filter` is used, it must return `ModelInputData` with a list input and the tracker must be updated with the filtered input (`mark_input_as_sent`). Session persistence is disabled when server-managed conversation is active.
|
||||
- Adding new tool/output/approval item types requires coordinated updates across:
|
||||
- `src/agents/items.py` (RunItem types and conversions)
|
||||
- `src/agents/run_internal/run_steps.py` (ProcessedResponse and tool run structs)
|
||||
- `src/agents/run_internal/turn_resolution.py` (model output processing, run item extraction)
|
||||
- `src/agents/run_internal/tool_execution.py` and `src/agents/run_internal/tool_planning.py`
|
||||
- `src/agents/run_internal/items.py` (normalization, dedupe, approval filtering)
|
||||
- `src/agents/stream_events.py` (stream event names)
|
||||
- `src/agents/run_state.py` (RunState serialization/deserialization)
|
||||
- `src/agents/run_internal/session_persistence.py` (session save/rewind)
|
||||
- If the serialized RunState shape changes, update `CURRENT_SCHEMA_VERSION` in `src/agents/run_state.py` and the related serialization/deserialization logic. Keep released schema versions readable, and feel free to renumber or squash unreleased schema versions before release when those intermediate snapshots are intentionally unsupported.
|
||||
- When bumping `CURRENT_SCHEMA_VERSION`, also add or update the matching entry in `SCHEMA_VERSION_SUMMARIES` in `src/agents/run_state.py` so every supported version keeps a short historical note describing what changed in that schema.
|
||||
|
||||
## Operation Guide
|
||||
|
||||
### Prerequisites
|
||||
|
||||
- Python 3.10+.
|
||||
- `uv` installed for dependency management (`uv sync`) and `uv run` for Python commands.
|
||||
- `make` available to run repository tasks.
|
||||
|
||||
### Development Workflow
|
||||
|
||||
1. Sync with `main` and create a feature branch:
|
||||
```bash
|
||||
git checkout -b feat/<short-description>
|
||||
```
|
||||
2. If dependencies changed or you are setting up the repo, run `make sync`.
|
||||
3. Implement changes and add or update tests alongside code updates.
|
||||
4. Highlight compatibility or API risks in your plan before implementing changes that alter the latest released behavior or a released or explicitly supported durable external state boundary.
|
||||
5. Build docs when you touch documentation:
|
||||
```bash
|
||||
make build-docs
|
||||
```
|
||||
6. When `$code-change-verification` applies, run it to execute the full verification stack before marking work complete.
|
||||
7. Commit with concise, imperative messages; keep commits small and focused, then open a pull request.
|
||||
8. When reporting code changes as complete (after substantial code work), invoke `$pr-draft-summary` as the final handoff step unless the task falls under the documented skip cases.
|
||||
|
||||
### Testing & Automated Checks
|
||||
|
||||
Before submitting changes, ensure relevant checks pass and extend tests when you touch code.
|
||||
|
||||
When `$code-change-verification` applies, run it to execute the required verification stack from the repository root. Rerun the full stack after applying fixes.
|
||||
|
||||
#### Unit tests and type checking
|
||||
|
||||
- Run the full test suite:
|
||||
```bash
|
||||
make tests
|
||||
```
|
||||
- Run a focused test:
|
||||
```bash
|
||||
uv run pytest -s -k <pattern>
|
||||
```
|
||||
- Type checking:
|
||||
```bash
|
||||
make typecheck
|
||||
```
|
||||
|
||||
#### Snapshot tests
|
||||
|
||||
Some tests rely on inline snapshots; see `tests/README.md` for details. Re-run `make tests` after updating snapshots.
|
||||
|
||||
- Fix snapshots:
|
||||
```bash
|
||||
make snapshots-fix
|
||||
```
|
||||
- Create new snapshots:
|
||||
```bash
|
||||
make snapshots-create
|
||||
```
|
||||
|
||||
#### Coverage
|
||||
|
||||
- Generate coverage (fails if coverage drops below threshold):
|
||||
```bash
|
||||
make coverage
|
||||
```
|
||||
|
||||
#### Formatting, linting, and type checking
|
||||
|
||||
- Formatting and linting use `ruff`; run `make format` (applies fixes) and `make lint` (checks only).
|
||||
- Type hints must pass `make typecheck`.
|
||||
- Write comments as full sentences ending with a period.
|
||||
- Imports are managed by Ruff and should stay sorted.
|
||||
|
||||
#### Mandatory local run order
|
||||
|
||||
When `$code-change-verification` applies, run the full sequence in order (or use the skill scripts):
|
||||
|
||||
```bash
|
||||
make format
|
||||
make lint
|
||||
make typecheck
|
||||
make tests
|
||||
```
|
||||
|
||||
### Utilities & Tips
|
||||
|
||||
- Install or refresh development dependencies:
|
||||
```bash
|
||||
make sync
|
||||
```
|
||||
- Run tests against the oldest supported version (Python 3.10) in an isolated environment:
|
||||
```bash
|
||||
UV_PROJECT_ENVIRONMENT=.venv_310 uv sync --python 3.10 --all-extras --all-packages --group dev
|
||||
UV_PROJECT_ENVIRONMENT=.venv_310 uv run --python 3.10 -m pytest
|
||||
```
|
||||
- Documentation workflows:
|
||||
```bash
|
||||
make build-docs # build docs after editing docs
|
||||
make serve-docs # preview docs locally
|
||||
make build-full-docs # run translations and build
|
||||
```
|
||||
- Snapshot helpers:
|
||||
```bash
|
||||
make snapshots-fix
|
||||
make snapshots-create
|
||||
```
|
||||
- Use `examples/` to see common SDK usage patterns.
|
||||
- Review `Makefile` for common commands and use `uv run` for Python invocations.
|
||||
- Explore `docs/` and `docs/scripts/` to understand the documentation pipeline.
|
||||
- Consult `tests/README.md` for test and snapshot workflows.
|
||||
- Check `mkdocs.yml` to understand how docs are organized.
|
||||
|
||||
### Pull Request & Commit Guidelines
|
||||
|
||||
- Use the template at `.github/PULL_REQUEST_TEMPLATE/pull_request_template.md`; include a summary, test plan, and issue number if applicable.
|
||||
- Add tests for new behavior when feasible and update documentation for user-facing changes.
|
||||
- Run `make format`, `make lint`, `make typecheck`, and `make tests` before marking work ready.
|
||||
- Commit messages should be concise and written in the imperative mood. Small, focused commits are preferred.
|
||||
|
||||
### Review Process & What Reviewers Look For
|
||||
|
||||
- ✅ Checks pass (`make format`, `make lint`, `make typecheck`, `make tests`).
|
||||
- ✅ Tests cover new behavior and edge cases.
|
||||
- ✅ Code is readable, maintainable, and consistent with existing style.
|
||||
- ✅ Public APIs and user-facing behavior changes are documented.
|
||||
- ✅ Examples are updated if behavior changes.
|
||||
- ✅ History is clean with a clear PR description.
|
||||
@@ -7,24 +7,56 @@ format:
|
||||
uv run ruff format
|
||||
uv run ruff check --fix
|
||||
|
||||
.PHONY: format-check
|
||||
format-check:
|
||||
uv run ruff format --check
|
||||
|
||||
.PHONY: lint
|
||||
lint:
|
||||
uv run ruff check
|
||||
|
||||
.PHONY: mypy
|
||||
mypy:
|
||||
uv run mypy .
|
||||
uv run mypy . --exclude site
|
||||
|
||||
.PHONY: pyright
|
||||
pyright:
|
||||
uv run pyright --project pyrightconfig.json
|
||||
|
||||
.PHONY: typecheck
|
||||
typecheck:
|
||||
@set -eu; \
|
||||
mypy_pid=''; \
|
||||
pyright_pid=''; \
|
||||
trap 'test -n "$$mypy_pid" && kill $$mypy_pid 2>/dev/null || true; test -n "$$pyright_pid" && kill $$pyright_pid 2>/dev/null || true' EXIT INT TERM; \
|
||||
echo "Running make mypy and make pyright in parallel..."; \
|
||||
$(MAKE) mypy & mypy_pid=$$!; \
|
||||
$(MAKE) pyright & pyright_pid=$$!; \
|
||||
wait $$mypy_pid; \
|
||||
wait $$pyright_pid; \
|
||||
trap - EXIT
|
||||
|
||||
.PHONY: tests
|
||||
tests:
|
||||
uv run pytest
|
||||
tests: tests-parallel tests-serial
|
||||
|
||||
.PHONY: tests-asyncio-stability
|
||||
tests-asyncio-stability:
|
||||
bash .github/scripts/run-asyncio-teardown-stability.sh
|
||||
|
||||
.PHONY: tests-parallel
|
||||
tests-parallel:
|
||||
uv run pytest -n auto --dist loadfile -m "not serial"
|
||||
|
||||
.PHONY: tests-serial
|
||||
tests-serial:
|
||||
uv run pytest -m serial
|
||||
|
||||
.PHONY: coverage
|
||||
coverage:
|
||||
|
||||
uv run coverage run -m pytest
|
||||
uv run coverage xml -o coverage.xml
|
||||
uv run coverage report -m --fail-under=95
|
||||
uv run coverage report -m --fail-under=85
|
||||
|
||||
.PHONY: snapshots-fix
|
||||
snapshots-fix:
|
||||
@@ -34,12 +66,9 @@ snapshots-fix:
|
||||
snapshots-create:
|
||||
uv run pytest --inline-snapshot=create
|
||||
|
||||
.PHONY: old_version_tests
|
||||
old_version_tests:
|
||||
UV_PROJECT_ENVIRONMENT=.venv_39 uv run --python 3.9 -m pytest
|
||||
|
||||
.PHONY: build-docs
|
||||
build-docs:
|
||||
uv run docs/scripts/generate_ref_files.py
|
||||
uv run mkdocs build
|
||||
|
||||
.PHONY: build-full-docs
|
||||
@@ -55,5 +84,5 @@ serve-docs:
|
||||
deploy-docs:
|
||||
uv run mkdocs gh-deploy --force --verbose
|
||||
|
||||
|
||||
|
||||
.PHONY: check
|
||||
check: format-check lint typecheck tests
|
||||
|
||||
@@ -0,0 +1,100 @@
|
||||
# Codex Execution Plans (ExecPlans)
|
||||
|
||||
This file defines how to write and maintain an ExecPlan: a self-contained, living specification that a novice can follow to deliver observable, working behavior in this repository.
|
||||
|
||||
## When to Use an ExecPlan
|
||||
- Required for multi-step or multi-file work, new features, refactors, or tasks expected to take more than about an hour.
|
||||
- Optional for trivial fixes (typos, small docs), but if you skip it for a substantial task, state the reason in your response.
|
||||
|
||||
## How to Use This File
|
||||
- Authoring: read this file end to end before drafting; start from the skeleton; embed all context (paths, commands, definitions) so no external docs are needed.
|
||||
- Implementing: move directly to the next milestone without asking for next steps; keep the living sections current at every stopping point.
|
||||
- Discussing: record decisions and rationale inside the plan so work can be resumed later using only the ExecPlan.
|
||||
|
||||
## Non-Negotiable Requirements
|
||||
- Self-contained and beginner-friendly: define every term; include needed repo knowledge; avoid assuming prior plans or external links.
|
||||
- Living document: revise Progress, Surprises & Discoveries, Decision Log, and Outcomes & Retrospective as work proceeds while keeping the plan self-contained.
|
||||
- Outcome-focused: describe what the user can do after the change and how to see it working; the plan must lead to demonstrably working behavior, not just code edits.
|
||||
- Explicit acceptance: state behaviors, commands, and observable outputs that prove success.
|
||||
|
||||
## Formatting Rules
|
||||
- Default envelope is a single fenced code block labeled `md`; do not nest other triple backticks inside—indent commands, transcripts, and diffs instead.
|
||||
- If the file contains only the ExecPlan, omit the enclosing code fence.
|
||||
- Use blank lines after headings; prefer prose over lists. Checklists are permitted only in the Progress section (and are mandatory there).
|
||||
|
||||
## Guidelines
|
||||
- Define jargon immediately and tie it to concrete files or commands in this repo.
|
||||
- Anchor on outcomes: acceptance should be phrased as observable behavior; for internal changes, show tests or scenarios that demonstrate the effect.
|
||||
- Specify repository context explicitly: full paths, functions, modules, working directory for commands, and environment assumptions.
|
||||
- Be idempotent and safe: describe retries or rollbacks for risky steps; prefer additive, testable changes.
|
||||
- Validation is required: state exact test commands and expected outputs; include concise evidence (logs, transcripts, diffs) as indented examples.
|
||||
|
||||
## Milestones
|
||||
- Tell a story (goal → work → result → proof) for each milestone; keep them narrative rather than bureaucratic.
|
||||
- Each milestone must be independently verifiable and incrementally advance the overall goal.
|
||||
- Milestones are distinct from Progress: milestones explain the plan; Progress tracks real-time execution.
|
||||
|
||||
## Living Sections (must be present and maintained)
|
||||
- Progress: checkbox list with timestamps; every pause should update what is done and what remains.
|
||||
- Surprises & Discoveries: unexpected behaviors, performance notes, or bugs with brief evidence.
|
||||
- Decision Log: each decision with rationale and date/author.
|
||||
- Outcomes & Retrospective: what was achieved, remaining gaps, and lessons learned.
|
||||
|
||||
## Prototyping and Parallel Paths
|
||||
- Prototypes are encouraged to de-risk changes; keep them additive, clearly labeled, and validated.
|
||||
- Parallel implementations are acceptable when reducing risk; describe how to validate each path and how to retire one safely.
|
||||
|
||||
## ExecPlan Skeleton
|
||||
|
||||
```md
|
||||
# <Short, action-oriented description>
|
||||
|
||||
This ExecPlan is a living document. The sections Progress, Surprises & Discoveries, Decision Log, and Outcomes & Retrospective must stay up to date as work proceeds.
|
||||
|
||||
If PLANS.md is present in the repo, maintain this document in accordance with it and link back to it by path.
|
||||
|
||||
## Purpose / Big Picture
|
||||
Explain the user-visible behavior gained after this change and how to observe it.
|
||||
|
||||
## Progress
|
||||
- [x] (2025-10-01 13:00Z) Example completed step.
|
||||
- [ ] Example incomplete step.
|
||||
- [ ] Example partially completed step (completed: X; remaining: Y).
|
||||
|
||||
## Surprises & Discoveries
|
||||
- Observation: …
|
||||
Evidence: …
|
||||
|
||||
## Decision Log
|
||||
- Decision: …
|
||||
Rationale: …
|
||||
Date/Author: …
|
||||
|
||||
## Outcomes & Retrospective
|
||||
Summarize outcomes, gaps, and lessons learned; compare to the original purpose.
|
||||
|
||||
## Context and Orientation
|
||||
Describe the current state relevant to this task as if the reader knows nothing. Name key files and modules by full path; define any non-obvious terms.
|
||||
|
||||
## Plan of Work
|
||||
Prose description of the sequence of edits and additions. For each edit, name the file and location and what to change.
|
||||
|
||||
## Concrete Steps
|
||||
Exact commands to run (with working directory). Include short expected outputs for comparison.
|
||||
|
||||
## Validation and Acceptance
|
||||
Behavioral acceptance criteria plus test commands and expected results.
|
||||
|
||||
## Idempotence and Recovery
|
||||
How to retry or roll back safely; ensure steps can be rerun without harm.
|
||||
|
||||
## Artifacts and Notes
|
||||
Concise transcripts, diffs, or snippets as indented examples.
|
||||
|
||||
## Interfaces and Dependencies
|
||||
Prescribe libraries, modules, and function signatures that must exist at the end. Use stable names and paths.
|
||||
```
|
||||
|
||||
## Revising a Plan
|
||||
- When the scope shifts, rewrite affected sections so the document remains coherent and self-contained.
|
||||
- After significant edits, add a short note at the end explaining what changed and why.
|
||||
@@ -1,180 +1,109 @@
|
||||
# OpenAI Agents SDK
|
||||
# OpenAI Agents SDK [](https://pypi.org/project/openai-agents/)
|
||||
|
||||
The OpenAI Agents SDK is a lightweight yet powerful framework for building multi-agent workflows.
|
||||
The OpenAI Agents SDK is a lightweight yet powerful framework for building multi-agent workflows. It is provider-agnostic, supporting the OpenAI Responses and Chat Completions APIs, as well as 100+ other LLMs.
|
||||
|
||||
<img src="https://cdn.openai.com/API/docs/images/orchestration.png" alt="Image of the Agents Tracing UI" style="max-height: 803px;">
|
||||
|
||||
> [!NOTE]
|
||||
> Looking for the JavaScript/TypeScript version? Check out [Agents SDK JS/TS](https://github.com/openai/openai-agents-js).
|
||||
|
||||
### Core concepts:
|
||||
|
||||
1. [**Agents**](https://openai.github.io/openai-agents-python/agents): LLMs configured with instructions, tools, guardrails, and handoffs
|
||||
2. [**Handoffs**](https://openai.github.io/openai-agents-python/handoffs/): A specialized tool call used by the Agents SDK for transferring control between agents
|
||||
3. [**Guardrails**](https://openai.github.io/openai-agents-python/guardrails/): Configurable safety checks for input and output validation
|
||||
4. [**Tracing**](https://openai.github.io/openai-agents-python/tracing/): Built-in tracking of agent runs, allowing you to view, debug and optimize your workflows
|
||||
1. [**Sandbox Agents**](https://openai.github.io/openai-agents-python/sandbox_agents): Agents preconfigured to work with a container to perform work over long time horizons.
|
||||
1. **[Agents as tools](https://openai.github.io/openai-agents-python/tools/#agents-as-tools) / [Handoffs](https://openai.github.io/openai-agents-python/handoffs/)**: Delegating to other agents for specific tasks
|
||||
1. [**Tools**](https://openai.github.io/openai-agents-python/tools/): Various Tools let agents take actions (functions, MCP, hosted tools)
|
||||
1. [**Guardrails**](https://openai.github.io/openai-agents-python/guardrails/): Configurable safety checks for input and output validation
|
||||
1. [**Human in the loop**](https://openai.github.io/openai-agents-python/human_in_the_loop/): Built-in mechanisms for involving humans across agent runs
|
||||
1. [**Sessions**](https://openai.github.io/openai-agents-python/sessions/): Automatic conversation history management across agent runs
|
||||
1. [**Tracing**](https://openai.github.io/openai-agents-python/tracing/): Built-in tracking of agent runs, allowing you to view, debug and optimize your workflows
|
||||
1. [**Realtime Agents**](https://openai.github.io/openai-agents-python/realtime/quickstart/): Build powerful voice agents with `gpt-realtime-2` and full agent features
|
||||
|
||||
Explore the [examples](examples) directory to see the SDK in action, and read our [documentation](https://openai.github.io/openai-agents-python/) for more details.
|
||||
|
||||
Notably, our SDK [is compatible](https://openai.github.io/openai-agents-python/models/) with any model providers that support the OpenAI Chat Completions API format.
|
||||
Explore the [examples](https://github.com/openai/openai-agents-python/tree/main/examples) directory to see the SDK in action, and read our [documentation](https://openai.github.io/openai-agents-python/) for more details.
|
||||
|
||||
## Get started
|
||||
|
||||
1. Set up your Python environment
|
||||
To get started, set up your Python environment (Python 3.10 or newer required), and then install OpenAI Agents SDK package.
|
||||
|
||||
```
|
||||
python -m venv env
|
||||
source env/bin/activate
|
||||
```
|
||||
### venv
|
||||
|
||||
2. Install Agents SDK
|
||||
|
||||
```
|
||||
```bash
|
||||
python -m venv .venv
|
||||
source .venv/bin/activate # On Windows: .venv\Scripts\activate
|
||||
pip install openai-agents
|
||||
```
|
||||
|
||||
For voice support, install with the optional `voice` group: `pip install 'openai-agents[voice]'`.
|
||||
For voice support, install with the optional `voice` group: `pip install 'openai-agents[voice]'`. For Redis session support, install with the optional `redis` group: `pip install 'openai-agents[redis]'`.
|
||||
|
||||
## Hello world example
|
||||
### uv
|
||||
|
||||
If you're familiar with [uv](https://docs.astral.sh/uv/), installing the package would be even easier:
|
||||
|
||||
```bash
|
||||
uv init
|
||||
uv add openai-agents
|
||||
```
|
||||
|
||||
For voice support, install with the optional `voice` group: `uv add 'openai-agents[voice]'`. For Redis session support, install with the optional `redis` group: `uv add 'openai-agents[redis]'`.
|
||||
|
||||
## Run your first Sandbox Agent
|
||||
|
||||
[Sandbox Agents](https://openai.github.io/openai-agents-python/sandbox_agents) are new in version 0.14.0. A sandbox agent is an agent that uses a computer environment to perform real work with a filesystem, in an environment you configure and control. Sandbox agents are useful when the agent needs to inspect files, run commands, apply patches, or carry workspace state across longer tasks.
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner
|
||||
from agents import Runner
|
||||
from agents.run import RunConfig
|
||||
from agents.sandbox import Manifest, SandboxAgent, SandboxRunConfig
|
||||
from agents.sandbox.entries import GitRepo
|
||||
from agents.sandbox.sandboxes import UnixLocalSandboxClient
|
||||
|
||||
agent = Agent(name="Assistant", instructions="You are a helpful assistant")
|
||||
agent = SandboxAgent(
|
||||
name="Workspace Assistant",
|
||||
instructions="Inspect the sandbox workspace before answering.",
|
||||
default_manifest=Manifest(
|
||||
entries={
|
||||
"repo": GitRepo(repo="openai/openai-agents-python", ref="main"),
|
||||
}
|
||||
),
|
||||
)
|
||||
|
||||
result = Runner.run_sync(agent, "Write a haiku about recursion in programming.")
|
||||
result = Runner.run_sync(
|
||||
agent,
|
||||
"Inspect the repo README and summarize what this project does.",
|
||||
# Run this agent on the local filesystem
|
||||
run_config=RunConfig(sandbox=SandboxRunConfig(client=UnixLocalSandboxClient())),
|
||||
)
|
||||
print(result.final_output)
|
||||
|
||||
# Code within the code,
|
||||
# Functions calling themselves,
|
||||
# Infinite loop's dance.
|
||||
# This project provides a Python SDK for building multi-agent workflows.
|
||||
```
|
||||
|
||||
(_If running this, ensure you set the `OPENAI_API_KEY` environment variable_)
|
||||
|
||||
(_For Jupyter notebook users, see [hello_world_jupyter.py](examples/basic/hello_world_jupyter.py)_)
|
||||
(_For Jupyter notebook users, see [hello_world_jupyter.ipynb](https://github.com/openai/openai-agents-python/blob/main/examples/basic/hello_world_jupyter.ipynb)_)
|
||||
|
||||
## Handoffs example
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner
|
||||
import asyncio
|
||||
|
||||
spanish_agent = Agent(
|
||||
name="Spanish agent",
|
||||
instructions="You only speak Spanish.",
|
||||
)
|
||||
|
||||
english_agent = Agent(
|
||||
name="English agent",
|
||||
instructions="You only speak English",
|
||||
)
|
||||
|
||||
triage_agent = Agent(
|
||||
name="Triage agent",
|
||||
instructions="Handoff to the appropriate agent based on the language of the request.",
|
||||
handoffs=[spanish_agent, english_agent],
|
||||
)
|
||||
|
||||
|
||||
async def main():
|
||||
result = await Runner.run(triage_agent, input="Hola, ¿cómo estás?")
|
||||
print(result.final_output)
|
||||
# ¡Hola! Estoy bien, gracias por preguntar. ¿Y tú, cómo estás?
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
## Functions example
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
|
||||
from agents import Agent, Runner, function_tool
|
||||
|
||||
|
||||
@function_tool
|
||||
def get_weather(city: str) -> str:
|
||||
return f"The weather in {city} is sunny."
|
||||
|
||||
|
||||
agent = Agent(
|
||||
name="Hello world",
|
||||
instructions="You are a helpful agent.",
|
||||
tools=[get_weather],
|
||||
)
|
||||
|
||||
|
||||
async def main():
|
||||
result = await Runner.run(agent, input="What's the weather in Tokyo?")
|
||||
print(result.final_output)
|
||||
# The weather in Tokyo is sunny.
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
## The agent loop
|
||||
|
||||
When you call `Runner.run()`, we run a loop until we get a final output.
|
||||
|
||||
1. We call the LLM, using the model and settings on the agent, and the message history.
|
||||
2. The LLM returns a response, which may include tool calls.
|
||||
3. If the response has a final output (see below for more on this), we return it and end the loop.
|
||||
4. If the response has a handoff, we set the agent to the new agent and go back to step 1.
|
||||
5. We process the tool calls (if any) and append the tool responses messages. Then we go to step 1.
|
||||
|
||||
There is a `max_turns` parameter that you can use to limit the number of times the loop executes.
|
||||
|
||||
### Final output
|
||||
|
||||
Final output is the last thing the agent produces in the loop.
|
||||
|
||||
1. If you set an `output_type` on the agent, the final output is when the LLM returns something of that type. We use [structured outputs](https://platform.openai.com/docs/guides/structured-outputs) for this.
|
||||
2. If there's no `output_type` (i.e. plain text responses), then the first LLM response without any tool calls or handoffs is considered as the final output.
|
||||
|
||||
As a result, the mental model for the agent loop is:
|
||||
|
||||
1. If the current agent has an `output_type`, the loop runs until the agent produces structured output matching that type.
|
||||
2. If the current agent does not have an `output_type`, the loop runs until the current agent produces a message without any tool calls/handoffs.
|
||||
|
||||
## Common agent patterns
|
||||
|
||||
The Agents SDK is designed to be highly flexible, allowing you to model a wide range of LLM workflows including deterministic flows, iterative loops, and more. See examples in [`examples/agent_patterns`](examples/agent_patterns).
|
||||
|
||||
## Tracing
|
||||
|
||||
The Agents SDK automatically traces your agent runs, making it easy to track and debug the behavior of your agents. Tracing is extensible by design, supporting custom spans and a wide variety of external destinations, including [Logfire](https://logfire.pydantic.dev/docs/integrations/llms/openai/#openai-agents), [AgentOps](https://docs.agentops.ai/v1/integrations/agentssdk), [Braintrust](https://braintrust.dev/docs/guides/traces/integrations#openai-agents-sdk), [Scorecard](https://docs.scorecard.io/docs/documentation/features/tracing#openai-agents-sdk-integration), and [Keywords AI](https://docs.keywordsai.co/integration/development-frameworks/openai-agent). For more details about how to customize or disable tracing, see [Tracing](http://openai.github.io/openai-agents-python/tracing), which also includes a larger list of [external tracing processors](http://openai.github.io/openai-agents-python/tracing/#external-tracing-processors-list).
|
||||
|
||||
## Development (only needed if you need to edit the SDK/examples)
|
||||
|
||||
0. Ensure you have [`uv`](https://docs.astral.sh/uv/) installed.
|
||||
|
||||
```bash
|
||||
uv --version
|
||||
```
|
||||
|
||||
1. Install dependencies
|
||||
|
||||
```bash
|
||||
make sync
|
||||
```
|
||||
|
||||
2. (After making changes) lint/test
|
||||
|
||||
```
|
||||
make tests # run tests
|
||||
make mypy # run typechecker
|
||||
make lint # run linter
|
||||
```
|
||||
Explore the [examples](https://github.com/openai/openai-agents-python/tree/main/examples) directory to see the SDK in action, and read our [documentation](https://openai.github.io/openai-agents-python/) for more details.
|
||||
|
||||
## Acknowledgements
|
||||
|
||||
We'd like to acknowledge the excellent work of the open-source community, especially:
|
||||
|
||||
- [Pydantic](https://docs.pydantic.dev/latest/) (data validation) and [PydanticAI](https://ai.pydantic.dev/) (advanced agent framework)
|
||||
- [MkDocs](https://github.com/squidfunk/mkdocs-material)
|
||||
- [Griffe](https://github.com/mkdocstrings/griffe)
|
||||
- [uv](https://github.com/astral-sh/uv) and [ruff](https://github.com/astral-sh/ruff)
|
||||
- [Pydantic](https://docs.pydantic.dev/latest/)
|
||||
- [Requests](https://github.com/psf/requests)
|
||||
- [MCP Python SDK](https://github.com/modelcontextprotocol/python-sdk)
|
||||
- [Griffe](https://github.com/mkdocstrings/griffe)
|
||||
|
||||
This library has these optional dependencies:
|
||||
|
||||
- [websockets](https://github.com/python-websockets/websockets)
|
||||
- [SQLAlchemy](https://github.com/sqlalchemy/sqlalchemy)
|
||||
- [any-llm](https://github.com/mozilla-ai/any-llm) and [LiteLLM](https://github.com/BerriAI/litellm)
|
||||
|
||||
We also rely on the following tools to manage the project:
|
||||
|
||||
- [uv](https://github.com/astral-sh/uv) and [ruff](https://github.com/astral-sh/ruff)
|
||||
- [mypy](https://github.com/python/mypy) and [Pyright](https://github.com/microsoft/pyright)
|
||||
- [pytest](https://github.com/pytest-dev/pytest) and [Coverage.py](https://github.com/coveragepy/coveragepy)
|
||||
- [MkDocs](https://github.com/squidfunk/mkdocs-material)
|
||||
|
||||
We're committed to continuing to build the Agents SDK as an open source framework so others in the community can expand on our approach.
|
||||
|
||||
@@ -0,0 +1,5 @@
|
||||
# Security Policy
|
||||
|
||||
For a more in-depth look at our security policy, please check out our [Coordinated Vulnerability Disclosure Policy](https://openai.com/security/disclosure/#:~:text=Disclosure%20Policy,-Security%20is%20essential&text=OpenAI%27s%20coordinated%20vulnerability%20disclosure%20policy,expect%20from%20us%20in%20return.).
|
||||
|
||||
Our PGP key can located [at this address.](https://cdn.openai.com/security.txt)
|
||||
+294
-16
@@ -1,37 +1,136 @@
|
||||
# Agents
|
||||
|
||||
Agents are the core building block in your apps. An agent is a large language model (LLM), configured with instructions and tools.
|
||||
Agents are the core building block in your apps. An agent is a large language model (LLM) configured with instructions, tools, and optional runtime behavior such as handoffs, guardrails, and structured outputs.
|
||||
|
||||
Use this page when you want to define or customize a single plain `Agent`. If you are deciding how multiple agents should collaborate, read [Agent orchestration](multi_agent.md). If the agent should run inside an isolated workspace with manifest-defined files and sandbox-native capabilities, read [Sandbox agent concepts](sandbox/guide.md).
|
||||
|
||||
The SDK uses the Responses API by default for OpenAI models, but the distinction here is orchestration: `Agent` plus `Runner` lets the SDK manage turns, tools, guardrails, handoffs, and sessions for you. If you want to own that loop yourself, use the Responses API directly instead.
|
||||
|
||||
## Choose the next guide
|
||||
|
||||
Use this page as the hub for agent definition. Jump to the adjacent guide that matches the next decision you need to make.
|
||||
|
||||
| If you want to... | Read next |
|
||||
| --- | --- |
|
||||
| Choose a model or provider setup | [Models](models/index.md) |
|
||||
| Add capabilities to the agent | [Tools](tools.md) |
|
||||
| Run an agent against a real repo, document bundle, or isolated workspace | [Sandbox agents quickstart](sandbox_agents.md) |
|
||||
| Decide between manager-style orchestration and handoffs | [Agent orchestration](multi_agent.md) |
|
||||
| Configure handoff behavior | [Handoffs](handoffs.md) |
|
||||
| Run turns, stream events, or manage conversation state | [Running agents](running_agents.md) |
|
||||
| Inspect final output, run items, or resumable state | [Results](results.md) |
|
||||
| Share local dependencies and runtime state | [Context management](context.md) |
|
||||
|
||||
## Basic configuration
|
||||
|
||||
The most common properties of an agent you'll configure are:
|
||||
The most common properties of an agent are:
|
||||
|
||||
- `instructions`: also known as a developer message or system prompt.
|
||||
- `model`: which LLM to use, and optional `model_settings` to configure model tuning parameters like temperature, top_p, etc.
|
||||
- `tools`: Tools that the agent can use to achieve its tasks.
|
||||
| Property | Required | Description |
|
||||
| --- | --- | --- |
|
||||
| `name` | yes | Human-readable agent name. |
|
||||
| `instructions` | no | System prompt or dynamic instructions callback. Strongly recommended. See [Dynamic instructions](#dynamic-instructions). |
|
||||
| `prompt` | no | OpenAI Responses API prompt configuration. Accepts a static prompt object or a function. See [Prompt templates](#prompt-templates). |
|
||||
| `handoff_description` | no | Short description exposed when this agent is offered as a handoff target. |
|
||||
| `handoffs` | no | Delegate the conversation to specialist agents. See [handoffs](handoffs.md). |
|
||||
| `model` | no | Which LLM to use. See [Models](models/index.md). |
|
||||
| `model_settings` | no | Model tuning parameters such as `temperature`, `top_p`, and `tool_choice`. |
|
||||
| `tools` | no | Tools the agent can call. See [Tools](tools.md). |
|
||||
| `mcp_servers` | no | MCP-backed tools for the agent. See the [MCP guide](mcp.md). |
|
||||
| `mcp_config` | no | Fine-tune how MCP tools are prepared, such as strict schema conversion and MCP failure formatting. See the [MCP guide](mcp.md#agent-level-mcp-configuration). |
|
||||
| `input_guardrails` | no | Guardrails that run on the first user input for this agent chain. See [Guardrails](guardrails.md). |
|
||||
| `output_guardrails` | no | Guardrails that run on the final output for this agent. See [Guardrails](guardrails.md). |
|
||||
| `output_type` | no | Structured output type instead of plain text. See [Output types](#output-types). |
|
||||
| `hooks` | no | Agent-scoped lifecycle callbacks. See [Lifecycle events (hooks)](#lifecycle-events-hooks). |
|
||||
| `tool_use_behavior` | no | Control whether tool results loop back to the model or end the run. See [Tool use behavior](#tool-use-behavior). |
|
||||
| `reset_tool_choice` | no | Reset `tool_choice` after a tool call (default: `True`) to avoid tool-use loops. See [Forcing tool use](#forcing-tool-use). |
|
||||
|
||||
```python
|
||||
from agents import Agent, ModelSettings, function_tool
|
||||
|
||||
@function_tool
|
||||
def get_weather(city: str) -> str:
|
||||
"""returns weather info for the specified city."""
|
||||
return f"The weather in {city} is sunny"
|
||||
|
||||
agent = Agent(
|
||||
name="Haiku agent",
|
||||
instructions="Always respond in haiku form",
|
||||
model="o3-mini",
|
||||
model="gpt-5-nano",
|
||||
tools=[get_weather],
|
||||
)
|
||||
```
|
||||
|
||||
Everything in this section applies to `Agent`. `SandboxAgent` builds on the same ideas, then adds `default_manifest`, `base_instructions`, `capabilities`, and `run_as` for workspace-scoped runs. See [Sandbox agent concepts](sandbox/guide.md).
|
||||
|
||||
## Prompt templates
|
||||
|
||||
You can reference a prompt template created in the OpenAI platform by setting `prompt`. This works with OpenAI models using the Responses API.
|
||||
|
||||
To use it, please:
|
||||
|
||||
1. Go to https://platform.openai.com/playground/prompts
|
||||
2. Create a new prompt variable, `poem_style`.
|
||||
3. Create a system prompt with the content:
|
||||
|
||||
```
|
||||
Write a poem in {{poem_style}}
|
||||
```
|
||||
|
||||
4. Run the example with the `--prompt-id` flag.
|
||||
|
||||
```python
|
||||
from agents import Agent
|
||||
|
||||
agent = Agent(
|
||||
name="Prompted assistant",
|
||||
prompt={
|
||||
"id": "pmpt_123",
|
||||
"version": "1",
|
||||
"variables": {"poem_style": "haiku"},
|
||||
},
|
||||
)
|
||||
```
|
||||
|
||||
You can also generate the prompt dynamically at run time:
|
||||
|
||||
```python
|
||||
from dataclasses import dataclass
|
||||
|
||||
from agents import Agent, GenerateDynamicPromptData, Runner
|
||||
|
||||
@dataclass
|
||||
class PromptContext:
|
||||
prompt_id: str
|
||||
poem_style: str
|
||||
|
||||
|
||||
async def build_prompt(data: GenerateDynamicPromptData):
|
||||
ctx: PromptContext = data.context.context
|
||||
return {
|
||||
"id": ctx.prompt_id,
|
||||
"version": "1",
|
||||
"variables": {"poem_style": ctx.poem_style},
|
||||
}
|
||||
|
||||
|
||||
agent = Agent(name="Prompted assistant", prompt=build_prompt)
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
"Say hello",
|
||||
context=PromptContext(prompt_id="pmpt_123", poem_style="limerick"),
|
||||
)
|
||||
```
|
||||
|
||||
## Context
|
||||
|
||||
Agents are generic on their `context` type. Context is a dependency-injection tool: it's an object you create and pass to `Runner.run()`, that is passed to every agent, tool, handoff etc, and it serves as a grab bag of dependencies and state for the agent run. You can provide any Python object as the context.
|
||||
|
||||
Read the [context guide](context.md) for the full `RunContextWrapper` surface, shared usage tracking, nested `tool_input`, and serialization caveats.
|
||||
|
||||
```python
|
||||
@dataclass
|
||||
class UserContext:
|
||||
name: str
|
||||
uid: str
|
||||
is_pro_user: bool
|
||||
|
||||
@@ -68,9 +167,47 @@ agent = Agent(
|
||||
|
||||
When you pass an `output_type`, that tells the model to use [structured outputs](https://platform.openai.com/docs/guides/structured-outputs) instead of regular plain text responses.
|
||||
|
||||
## Handoffs
|
||||
## Multi-agent system design patterns
|
||||
|
||||
Handoffs are sub-agents that the agent can delegate to. You provide a list of handoffs, and the agent can choose to delegate to them if relevant. This is a powerful pattern that allows orchestrating modular, specialized agents that excel at a single task. Read more in the [handoffs](handoffs.md) documentation.
|
||||
There are many ways to design multi‑agent systems, but we commonly see two broadly applicable patterns:
|
||||
|
||||
1. Manager (agents as tools): A central manager/orchestrator invokes specialized sub‑agents as tools and retains control of the conversation.
|
||||
2. Handoffs: Peer agents hand off control to a specialized agent that takes over the conversation. This is decentralized.
|
||||
|
||||
See [our practical guide to building agents](https://cdn.openai.com/business-guides-and-resources/a-practical-guide-to-building-agents.pdf) for more details.
|
||||
|
||||
### Manager (agents as tools)
|
||||
|
||||
The `customer_facing_agent` handles all user interaction and invokes specialized sub‑agents exposed as tools. Read more in the [tools](tools.md#agents-as-tools) documentation.
|
||||
|
||||
```python
|
||||
from agents import Agent
|
||||
|
||||
booking_agent = Agent(...)
|
||||
refund_agent = Agent(...)
|
||||
|
||||
customer_facing_agent = Agent(
|
||||
name="Customer-facing agent",
|
||||
instructions=(
|
||||
"Handle all direct user communication. "
|
||||
"Call the relevant tools when specialized expertise is needed."
|
||||
),
|
||||
tools=[
|
||||
booking_agent.as_tool(
|
||||
tool_name="booking_expert",
|
||||
tool_description="Handles booking questions and requests.",
|
||||
),
|
||||
refund_agent.as_tool(
|
||||
tool_name="refund_expert",
|
||||
tool_description="Handles refund questions and requests.",
|
||||
)
|
||||
],
|
||||
)
|
||||
```
|
||||
|
||||
### Handoffs
|
||||
|
||||
Handoffs are sub‑agents the agent can delegate to. When a handoff occurs, the delegated agent receives the conversation history and takes over the conversation. This pattern enables modular, specialized agents that excel at a single task. Read more in the [handoffs](handoffs.md) documentation.
|
||||
|
||||
```python
|
||||
from agents import Agent
|
||||
@@ -81,9 +218,9 @@ refund_agent = Agent(...)
|
||||
triage_agent = Agent(
|
||||
name="Triage agent",
|
||||
instructions=(
|
||||
"Help the user with their questions."
|
||||
"If they ask about booking, handoff to the booking agent."
|
||||
"If they ask about refunds, handoff to the refund agent."
|
||||
"Help the user with their questions. "
|
||||
"If they ask about booking, hand off to the booking agent. "
|
||||
"If they ask about refunds, hand off to the refund agent."
|
||||
),
|
||||
handoffs=[booking_agent, refund_agent],
|
||||
)
|
||||
@@ -108,11 +245,53 @@ agent = Agent[UserContext](
|
||||
|
||||
## Lifecycle events (hooks)
|
||||
|
||||
Sometimes, you want to observe the lifecycle of an agent. For example, you may want to log events, or pre-fetch data when certain events occur. You can hook into the agent lifecycle with the `hooks` property. Subclass the [`AgentHooks`][agents.lifecycle.AgentHooks] class, and override the methods you're interested in.
|
||||
Sometimes, you want to observe the lifecycle of an agent. For example, you may want to log events, pre-fetch data, or record usage when certain events occur.
|
||||
|
||||
There are two hook scopes:
|
||||
|
||||
- [`RunHooks`][agents.lifecycle.RunHooks] observe the entire `Runner.run(...)` invocation, including handoffs to other agents.
|
||||
- [`AgentHooks`][agents.lifecycle.AgentHooks] are attached to a specific agent instance via `agent.hooks`.
|
||||
|
||||
The callback context also changes depending on the event:
|
||||
|
||||
- Agent start/end hooks receive [`AgentHookContext`][agents.run_context.AgentHookContext], which wraps your original context and carries the shared run usage state.
|
||||
- LLM, tool, and handoff hooks receive [`RunContextWrapper`][agents.run_context.RunContextWrapper].
|
||||
|
||||
Typical hook timing:
|
||||
|
||||
- `on_agent_start` / `on_agent_end`: when a specific agent begins or finishes producing a final output.
|
||||
- `on_llm_start` / `on_llm_end`: immediately around each model call.
|
||||
- `on_tool_start` / `on_tool_end`: around each local tool invocation.
|
||||
For function tools, the hook `context` is typically a `ToolContext`, so you can inspect tool-call metadata such as `tool_call_id`.
|
||||
- `on_handoff`: when control moves from one agent to another.
|
||||
|
||||
Use `RunHooks` when you want a single observer for the whole workflow, and `AgentHooks` when one agent needs custom side effects.
|
||||
|
||||
```python
|
||||
from agents import Agent, RunHooks, Runner
|
||||
|
||||
|
||||
class LoggingHooks(RunHooks):
|
||||
async def on_agent_start(self, context, agent):
|
||||
print(f"Starting {agent.name}")
|
||||
|
||||
async def on_llm_end(self, context, agent, response):
|
||||
print(f"{agent.name} produced {len(response.output)} output items")
|
||||
|
||||
async def on_agent_end(self, context, agent, output):
|
||||
print(f"{agent.name} finished with usage: {context.usage}")
|
||||
|
||||
|
||||
agent = Agent(name="Assistant", instructions="Be concise.")
|
||||
result = await Runner.run(agent, "Explain quines", hooks=LoggingHooks())
|
||||
print(result.final_output)
|
||||
```
|
||||
|
||||
For the full callback surface, see the [Lifecycle API reference](ref/lifecycle.md).
|
||||
|
||||
## Guardrails
|
||||
|
||||
Guardrails allow you to run checks/validations on user input, in parallel to the agent running. For example, you could screen the user's input for relevance. Read more in the [guardrails](guardrails.md) documentation.
|
||||
Guardrails allow you to run checks/validations on user input in parallel to the agent running, and on the agent's output once it is produced. For example, you could screen the user's input and agent's output for relevance. Read more in the [guardrails](guardrails.md) documentation.
|
||||
|
||||
## Cloning/copying agents
|
||||
|
||||
@@ -122,7 +301,7 @@ By using the `clone()` method on an agent, you can duplicate an Agent, and optio
|
||||
pirate_agent = Agent(
|
||||
name="Pirate",
|
||||
instructions="Write like a pirate",
|
||||
model="o3-mini",
|
||||
model="gpt-5.5",
|
||||
)
|
||||
|
||||
robot_agent = pirate_agent.clone(
|
||||
@@ -140,8 +319,107 @@ Supplying a list of tools doesn't always mean the LLM will use a tool. You can f
|
||||
3. `none`, which requires the LLM to _not_ use a tool.
|
||||
4. Setting a specific string e.g. `my_tool`, which requires the LLM to use that specific tool.
|
||||
|
||||
When you are using OpenAI Responses tool search, named tool choices are more limited: you cannot target bare namespace names or deferred-only tools with `tool_choice`, and `tool_choice="tool_search"` does not target [`ToolSearchTool`][agents.tool.ToolSearchTool]. In those cases, prefer `auto` or `required`. See [Hosted tool search](tools.md#hosted-tool-search) for the Responses-specific constraints.
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner, function_tool, ModelSettings
|
||||
|
||||
@function_tool
|
||||
def get_weather(city: str) -> str:
|
||||
"""Returns weather info for the specified city."""
|
||||
return f"The weather in {city} is sunny"
|
||||
|
||||
agent = Agent(
|
||||
name="Weather Agent",
|
||||
instructions="Retrieve weather details.",
|
||||
tools=[get_weather],
|
||||
model_settings=ModelSettings(tool_choice="get_weather")
|
||||
)
|
||||
```
|
||||
|
||||
## Tool use behavior
|
||||
|
||||
The `tool_use_behavior` parameter in the `Agent` configuration controls how tool outputs are handled:
|
||||
|
||||
- `"run_llm_again"`: The default. Tools are run, and the LLM processes the results to produce a final response.
|
||||
- `"stop_on_first_tool"`: The output of the first tool call is used as the final response, without further LLM processing.
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner, function_tool, ModelSettings
|
||||
|
||||
@function_tool
|
||||
def get_weather(city: str) -> str:
|
||||
"""Returns weather info for the specified city."""
|
||||
return f"The weather in {city} is sunny"
|
||||
|
||||
agent = Agent(
|
||||
name="Weather Agent",
|
||||
instructions="Retrieve weather details.",
|
||||
tools=[get_weather],
|
||||
tool_use_behavior="stop_on_first_tool"
|
||||
)
|
||||
```
|
||||
|
||||
- `StopAtTools(stop_at_tool_names=[...])`: Stops if any specified tool is called, using its output as the final response.
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner, function_tool
|
||||
from agents.agent import StopAtTools
|
||||
|
||||
@function_tool
|
||||
def get_weather(city: str) -> str:
|
||||
"""Returns weather info for the specified city."""
|
||||
return f"The weather in {city} is sunny"
|
||||
|
||||
@function_tool
|
||||
def sum_numbers(a: int, b: int) -> int:
|
||||
"""Adds two numbers."""
|
||||
return a + b
|
||||
|
||||
agent = Agent(
|
||||
name="Stop At Stock Agent",
|
||||
instructions="Get weather or sum numbers.",
|
||||
tools=[get_weather, sum_numbers],
|
||||
tool_use_behavior=StopAtTools(stop_at_tool_names=["get_weather"])
|
||||
)
|
||||
```
|
||||
|
||||
- `ToolsToFinalOutputFunction`: A custom function that processes tool results and decides whether to stop or continue with the LLM.
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner, function_tool, FunctionToolResult, RunContextWrapper
|
||||
from agents.agent import ToolsToFinalOutputResult
|
||||
from typing import List, Any
|
||||
|
||||
@function_tool
|
||||
def get_weather(city: str) -> str:
|
||||
"""Returns weather info for the specified city."""
|
||||
return f"The weather in {city} is sunny"
|
||||
|
||||
def custom_tool_handler(
|
||||
context: RunContextWrapper[Any],
|
||||
tool_results: List[FunctionToolResult]
|
||||
) -> ToolsToFinalOutputResult:
|
||||
"""Processes tool results to decide final output."""
|
||||
for result in tool_results:
|
||||
if result.output and "sunny" in result.output:
|
||||
return ToolsToFinalOutputResult(
|
||||
is_final_output=True,
|
||||
final_output=f"Final weather: {result.output}"
|
||||
)
|
||||
return ToolsToFinalOutputResult(
|
||||
is_final_output=False,
|
||||
final_output=None
|
||||
)
|
||||
|
||||
agent = Agent(
|
||||
name="Weather Agent",
|
||||
instructions="Retrieve weather details.",
|
||||
tools=[get_weather],
|
||||
tool_use_behavior=custom_tool_handler
|
||||
)
|
||||
```
|
||||
|
||||
!!! note
|
||||
|
||||
To prevent infinite loops, the framework automatically resets `tool_choice` to "auto" after a tool call. This behavior is configurable via [`agent.reset_tool_choice`][agents.agent.Agent.reset_tool_choice]. The infinite loop is because tool results are sent to the LLM, which then generates another tool call because of `tool_choice`, ad infinitum.
|
||||
|
||||
If you want the Agent to completely stop after a tool call (rather than continuing with auto mode), you can set [`Agent.tool_use_behavior="stop_on_first_tool"`] which will directly use the tool output as the final response without further LLM processing.
|
||||
|
||||
Binary file not shown.
|
Before Width: | Height: | Size: 93 KiB After Width: | Height: | Size: 34 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 84 KiB |
+84
-9
@@ -1,8 +1,20 @@
|
||||
# Configuring the SDK
|
||||
# Configuration
|
||||
|
||||
This page covers SDK-wide defaults that you usually set once during application startup, such as the default OpenAI key or client, the default OpenAI API shape, tracing export defaults, and logging behavior.
|
||||
|
||||
These defaults still apply to sandbox-based workflows, but sandbox workspaces, sandbox clients, and session reuse are configured separately.
|
||||
|
||||
If you need to configure a specific agent or run instead, start with:
|
||||
|
||||
- [Agents](agents.md) for instructions, tools, output types, handoffs, and guardrails on a plain `Agent`.
|
||||
- [Running agents](running_agents.md) for `RunConfig`, sessions, and conversation-state options.
|
||||
- [Sandbox agents](sandbox/guide.md) for `SandboxRunConfig`, manifests, capabilities, and sandbox-client-specific workspace setup.
|
||||
- [Models](models/index.md) for model selection and provider configuration.
|
||||
- [Tracing](tracing.md) for per-run tracing metadata and custom trace processors.
|
||||
|
||||
## API keys and clients
|
||||
|
||||
By default, the SDK looks for the `OPENAI_API_KEY` environment variable for LLM requests and tracing, as soon as it is imported. If you are unable to set that environment variable before your app starts, you can use the [set_default_openai_key()][agents.set_default_openai_key] function to set the key.
|
||||
By default, the SDK uses the `OPENAI_API_KEY` environment variable for LLM requests and tracing. The key is resolved when the SDK first creates an OpenAI client (lazy initialization), so set the environment variable before your first model call. If you are unable to set that environment variable before your app starts, you can use the [set_default_openai_key()][agents.set_default_openai_key] function to set the key.
|
||||
|
||||
```python
|
||||
from agents import set_default_openai_key
|
||||
@@ -20,6 +32,13 @@ custom_client = AsyncOpenAI(base_url="...", api_key="...")
|
||||
set_default_openai_client(custom_client)
|
||||
```
|
||||
|
||||
If you prefer environment-based endpoint configuration, the default OpenAI provider also reads `OPENAI_BASE_URL`. When you enable Responses websocket transport, it also reads `OPENAI_WEBSOCKET_BASE_URL` for the websocket `/responses` endpoint.
|
||||
|
||||
```bash
|
||||
export OPENAI_BASE_URL="https://your-openai-compatible-endpoint.example/v1"
|
||||
export OPENAI_WEBSOCKET_BASE_URL="wss://your-openai-compatible-endpoint.example/v1"
|
||||
```
|
||||
|
||||
Finally, you can also customize the OpenAI API that is used. By default, we use the OpenAI Responses API. You can override this to use the Chat Completions API by using the [set_default_openai_api()][agents.set_default_openai_api] function.
|
||||
|
||||
```python
|
||||
@@ -30,7 +49,7 @@ set_default_openai_api("chat_completions")
|
||||
|
||||
## Tracing
|
||||
|
||||
Tracing is enabled by default. It uses the OpenAI API keys from the section above by default (i.e. the environment variable or the default key you set). You can specifically set the API key used for tracing by using the [`set_tracing_export_api_key`][agents.set_tracing_export_api_key] function.
|
||||
Tracing is enabled by default. By default it uses the same OpenAI API key as your model requests from the section above (that is, the environment variable or the default key you set). You can specifically set the API key used for tracing by using the [`set_tracing_export_api_key`][agents.set_tracing_export_api_key] function.
|
||||
|
||||
```python
|
||||
from agents import set_tracing_export_api_key
|
||||
@@ -38,6 +57,40 @@ from agents import set_tracing_export_api_key
|
||||
set_tracing_export_api_key("sk-...")
|
||||
```
|
||||
|
||||
If your model traffic uses one key or client but tracing should use a different OpenAI key, pass `use_for_tracing=False` when setting the default key or client, then configure tracing separately. The same pattern works with [`set_default_openai_key()`][agents.set_default_openai_key] if you are not using a custom client.
|
||||
|
||||
```python
|
||||
from openai import AsyncOpenAI
|
||||
from agents import (
|
||||
set_default_openai_client,
|
||||
set_tracing_export_api_key,
|
||||
)
|
||||
|
||||
custom_client = AsyncOpenAI(base_url="https://your-openai-compatible-endpoint.example/v1", api_key="provider-key")
|
||||
set_default_openai_client(custom_client, use_for_tracing=False)
|
||||
|
||||
set_tracing_export_api_key("sk-tracing")
|
||||
```
|
||||
|
||||
If you need to attribute traces to a specific organization or project when using the default exporter, set these environment variables before your app starts:
|
||||
|
||||
```bash
|
||||
export OPENAI_ORG_ID="org_..."
|
||||
export OPENAI_PROJECT_ID="proj_..."
|
||||
```
|
||||
|
||||
You can also set a tracing API key per run without changing the global exporter.
|
||||
|
||||
```python
|
||||
from agents import Runner, RunConfig
|
||||
|
||||
await Runner.run(
|
||||
agent,
|
||||
input="Hello",
|
||||
run_config=RunConfig(tracing={"api_key": "sk-tracing-123"}),
|
||||
)
|
||||
```
|
||||
|
||||
You can also disable tracing entirely by using the [`set_tracing_disabled()`][agents.set_tracing_disabled] function.
|
||||
|
||||
```python
|
||||
@@ -46,9 +99,29 @@ from agents import set_tracing_disabled
|
||||
set_tracing_disabled(True)
|
||||
```
|
||||
|
||||
If you want to keep tracing enabled but exclude potentially sensitive inputs/outputs from trace payloads, set [`RunConfig.trace_include_sensitive_data`][agents.run.RunConfig.trace_include_sensitive_data] to `False`:
|
||||
|
||||
```python
|
||||
from agents import Runner, RunConfig
|
||||
|
||||
await Runner.run(
|
||||
agent,
|
||||
input="Hello",
|
||||
run_config=RunConfig(trace_include_sensitive_data=False),
|
||||
)
|
||||
```
|
||||
|
||||
You can also change the default without code by setting this environment variable before your app starts:
|
||||
|
||||
```bash
|
||||
export OPENAI_AGENTS_TRACE_INCLUDE_SENSITIVE_DATA=0
|
||||
```
|
||||
|
||||
For full tracing controls, see the [tracing guide](tracing.md).
|
||||
|
||||
## Debug logging
|
||||
|
||||
The SDK has two Python loggers without any handlers set. By default, this means that warnings and errors are sent to `stdout`, but other logs are suppressed.
|
||||
The SDK defines two Python loggers (`openai.agents` and `openai.agents.tracing`) and does not attach handlers by default. Logs follow your application's Python logging configuration.
|
||||
|
||||
To enable verbose logging, use the [`enable_verbose_stdout_logging()`][agents.enable_verbose_stdout_logging] function.
|
||||
|
||||
@@ -79,16 +152,18 @@ logger.addHandler(logging.StreamHandler())
|
||||
|
||||
### Sensitive data in logs
|
||||
|
||||
Certain logs may contain sensitive data (for example, user data). If you want to disable this data from being logged, set the following environment variables.
|
||||
Certain logs may contain sensitive data (for example, user data).
|
||||
|
||||
To disable logging LLM inputs and outputs:
|
||||
By default, the SDK does **not** log LLM inputs/outputs or tool inputs/outputs. These protections are controlled by:
|
||||
|
||||
```bash
|
||||
export OPENAI_AGENTS_DONT_LOG_MODEL_DATA=1
|
||||
OPENAI_AGENTS_DONT_LOG_MODEL_DATA=1
|
||||
OPENAI_AGENTS_DONT_LOG_TOOL_DATA=1
|
||||
```
|
||||
|
||||
To disable logging tool inputs and outputs:
|
||||
If you need to include this data temporarily for debugging, set either variable to `0` (or `false`) before your app starts:
|
||||
|
||||
```bash
|
||||
export OPENAI_AGENTS_DONT_LOG_TOOL_DATA=1
|
||||
export OPENAI_AGENTS_DONT_LOG_MODEL_DATA=0
|
||||
export OPENAI_AGENTS_DONT_LOG_TOOL_DATA=0
|
||||
```
|
||||
|
||||
+69
-2
@@ -10,9 +10,11 @@ Context is an overloaded term. There are two main classes of context you might c
|
||||
This is represented via the [`RunContextWrapper`][agents.run_context.RunContextWrapper] class and the [`context`][agents.run_context.RunContextWrapper.context] property within it. The way this works is:
|
||||
|
||||
1. You create any Python object you want. A common pattern is to use a dataclass or a Pydantic object.
|
||||
2. You pass that object to the various run methods (e.g. `Runner.run(..., **context=whatever**))`.
|
||||
2. You pass that object to the various run methods (e.g. `Runner.run(..., context=whatever)`).
|
||||
3. All your tool calls, lifecycle hooks etc will be passed a wrapper object, `RunContextWrapper[T]`, where `T` represents your context object type which you can access via `wrapper.context`.
|
||||
|
||||
For some runtime-specific callbacks, the SDK may pass a more specialized subclass of `RunContextWrapper[T]`. For example, function-tool lifecycle hooks typically receive `ToolContext`, which also exposes tool-call metadata like `tool_call_id`, `tool_name`, and `tool_arguments`.
|
||||
|
||||
The **most important** thing to be aware of: every agent, tool function, lifecycle etc for a given agent run must use the same _type_ of context.
|
||||
|
||||
You can use the context for things like:
|
||||
@@ -25,6 +27,23 @@ You can use the context for things like:
|
||||
|
||||
The context object is **not** sent to the LLM. It is purely a local object that you can read from, write to and call methods on it.
|
||||
|
||||
Within a single run, derived wrappers share the same underlying app context, approval state, and usage tracking. Nested [`Agent.as_tool()`][agents.agent.Agent.as_tool] runs may attach a different `tool_input`, but they do not get an isolated copy of your app state by default.
|
||||
|
||||
### What `RunContextWrapper` exposes
|
||||
|
||||
[`RunContextWrapper`][agents.run_context.RunContextWrapper] is a wrapper around your app-defined context object. In practice you will most often use:
|
||||
|
||||
- [`wrapper.context`][agents.run_context.RunContextWrapper.context] for your own mutable app state and dependencies.
|
||||
- [`wrapper.usage`][agents.run_context.RunContextWrapper.usage] for aggregated request and token usage across the current run.
|
||||
- [`wrapper.tool_input`][agents.run_context.RunContextWrapper.tool_input] for structured input when the current run is executing inside [`Agent.as_tool()`][agents.agent.Agent.as_tool].
|
||||
- [`wrapper.approve_tool(...)`][agents.run_context.RunContextWrapper.approve_tool] / [`wrapper.reject_tool(...)`][agents.run_context.RunContextWrapper.reject_tool] when you need to update approval state programmatically.
|
||||
|
||||
Only `wrapper.context` is your app-defined object. The other fields are runtime metadata managed by the SDK.
|
||||
|
||||
If you later serialize a [`RunState`][agents.run_state.RunState] for human-in-the-loop or durable job workflows, that runtime metadata is saved with the state. Avoid putting secrets in [`RunContextWrapper.context`][agents.run_context.RunContextWrapper.context] if you intend to persist or transmit serialized state.
|
||||
|
||||
Conversation state is a separate concern. Use `result.to_input_list()`, `session`, `conversation_id`, or `previous_response_id` depending on how you want to carry turns forward. See [results](results.md), [running agents](running_agents.md), and [sessions](sessions/index.md) for that decision.
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
from dataclasses import dataclass
|
||||
@@ -38,7 +57,8 @@ class UserInfo: # (1)!
|
||||
|
||||
@function_tool
|
||||
async def fetch_user_age(wrapper: RunContextWrapper[UserInfo]) -> str: # (2)!
|
||||
return f"User {wrapper.context.name} is 47 years old"
|
||||
"""Fetch the age of the user. Call this function to get user's age information."""
|
||||
return f"The user {wrapper.context.name} is 47 years old"
|
||||
|
||||
async def main():
|
||||
user_info = UserInfo(name="John", uid=123)
|
||||
@@ -67,6 +87,53 @@ if __name__ == "__main__":
|
||||
4. The context is passed to the `run` function.
|
||||
5. The agent correctly calls the tool and gets the age.
|
||||
|
||||
---
|
||||
|
||||
### Advanced: `ToolContext`
|
||||
|
||||
In some cases, you might want to access extra metadata about the tool being executed — such as its name, call ID, or raw argument string.
|
||||
For this, you can use the [`ToolContext`][agents.tool_context.ToolContext] class, which extends `RunContextWrapper`.
|
||||
|
||||
```python
|
||||
from typing import Annotated
|
||||
from pydantic import BaseModel, Field
|
||||
from agents import Agent, Runner, function_tool
|
||||
from agents.tool_context import ToolContext
|
||||
|
||||
class WeatherContext(BaseModel):
|
||||
user_id: str
|
||||
|
||||
class Weather(BaseModel):
|
||||
city: str = Field(description="The city name")
|
||||
temperature_range: str = Field(description="The temperature range in Celsius")
|
||||
conditions: str = Field(description="The weather conditions")
|
||||
|
||||
@function_tool
|
||||
def get_weather(ctx: ToolContext[WeatherContext], city: Annotated[str, "The city to get the weather for"]) -> Weather:
|
||||
print(f"[debug] Tool context: (name: {ctx.tool_name}, call_id: {ctx.tool_call_id}, args: {ctx.tool_arguments})")
|
||||
return Weather(city=city, temperature_range="14-20C", conditions="Sunny with wind.")
|
||||
|
||||
agent = Agent(
|
||||
name="Weather Agent",
|
||||
instructions="You are a helpful agent that can tell the weather of a given city.",
|
||||
tools=[get_weather],
|
||||
)
|
||||
```
|
||||
|
||||
`ToolContext` provides the same `.context` property as `RunContextWrapper`,
|
||||
plus additional fields specific to the current tool call:
|
||||
|
||||
- `tool_name` – the name of the tool being invoked
|
||||
- `tool_call_id` – a unique identifier for this tool call
|
||||
- `tool_arguments` – the raw argument string passed to the tool
|
||||
- `tool_namespace` – the Responses namespace for the tool call, when the tool was loaded through `tool_namespace()` or another namespaced surface
|
||||
- `qualified_tool_name` – the tool name qualified with the namespace when one is available
|
||||
|
||||
Use `ToolContext` when you need tool-level metadata during execution.
|
||||
For general context sharing between agents and tools, `RunContextWrapper` remains sufficient. Because `ToolContext` extends `RunContextWrapper`, it can also expose `.tool_input` when a nested `Agent.as_tool()` run supplied structured input.
|
||||
|
||||
---
|
||||
|
||||
## Agent/LLM context
|
||||
|
||||
When an LLM is called, the **only** data it can see is from the conversation history. This means that if you want to make some new data available to the LLM, you must do it in a way that makes it available in that history. There are a few ways to do this:
|
||||
|
||||
+122
-26
@@ -2,41 +2,137 @@
|
||||
|
||||
Check out a variety of sample implementations of the SDK in the examples section of the [repo](https://github.com/openai/openai-agents-python/tree/main/examples). The examples are organized into several categories that demonstrate different patterns and capabilities.
|
||||
|
||||
|
||||
## Categories
|
||||
|
||||
- **[agent_patterns](https://github.com/openai/openai-agents-python/tree/main/examples/agent_patterns):**
|
||||
Examples in this category illustrate common agent design patterns, such as
|
||||
- **[agent_patterns](https://github.com/openai/openai-agents-python/tree/main/examples/agent_patterns):**
|
||||
Examples in this category illustrate common agent design patterns, such as
|
||||
|
||||
- Deterministic workflows
|
||||
- Agents as tools
|
||||
- Parallel agent execution
|
||||
- Deterministic workflows
|
||||
- Agents as tools
|
||||
- Agents as tools with streaming events (`examples/agent_patterns/agents_as_tools_streaming.py`)
|
||||
- Agents as tools with structured input parameters (`examples/agent_patterns/agents_as_tools_structured.py`)
|
||||
- Parallel agent execution
|
||||
- Conditional tool usage
|
||||
- Forcing tool use with different behaviors (`examples/agent_patterns/forcing_tool_use.py`)
|
||||
- Input/output guardrails
|
||||
- LLM as a judge
|
||||
- Routing
|
||||
- Streaming guardrails
|
||||
- Human-in-the-loop with tool approval and state serialization (`examples/agent_patterns/human_in_the_loop.py`)
|
||||
- Human-in-the-loop with streaming (`examples/agent_patterns/human_in_the_loop_stream.py`)
|
||||
- Custom rejection messages for approval flows (`examples/agent_patterns/human_in_the_loop_custom_rejection.py`)
|
||||
|
||||
- **[basic](https://github.com/openai/openai-agents-python/tree/main/examples/basic):**
|
||||
These examples showcase foundational capabilities of the SDK, such as
|
||||
- **[basic](https://github.com/openai/openai-agents-python/tree/main/examples/basic):**
|
||||
These examples showcase foundational capabilities of the SDK, such as
|
||||
|
||||
- Dynamic system prompts
|
||||
- Streaming outputs
|
||||
- Lifecycle events
|
||||
- Hello world examples (Default model, GPT-5, open-weight model)
|
||||
- Agent lifecycle management
|
||||
- Run hooks and agent hooks lifecycle example (`examples/basic/lifecycle_example.py`)
|
||||
- Dynamic system prompts
|
||||
- Basic tool usage (`examples/basic/tools.py`)
|
||||
- Tool input/output guardrails (`examples/basic/tool_guardrails.py`)
|
||||
- Image tool output (`examples/basic/image_tool_output.py`)
|
||||
- Streaming outputs (text, items, function call args)
|
||||
- Responses websocket transport with a shared session helper across turns (`examples/basic/stream_ws.py`)
|
||||
- Prompt templates
|
||||
- File handling (local and remote, images and PDFs)
|
||||
- Usage tracking
|
||||
- Runner-managed retry settings (`examples/basic/retry.py`)
|
||||
- Runner-managed retries through a third-party adapter (`examples/basic/retry_litellm.py`)
|
||||
- Non-strict output types
|
||||
- Previous response ID usage
|
||||
|
||||
- **[tool examples](https://github.com/openai/openai-agents-python/tree/main/examples/tools):**
|
||||
Learn how to implement OAI hosted tools such as web search and file search,
|
||||
and integrate them into your agents.
|
||||
- **[customer_service](https://github.com/openai/openai-agents-python/tree/main/examples/customer_service):**
|
||||
Example customer service system for an airline.
|
||||
|
||||
- **[model providers](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers):**
|
||||
Explore how to use non-OpenAI models with the SDK.
|
||||
- **[financial_research_agent](https://github.com/openai/openai-agents-python/tree/main/examples/financial_research_agent):**
|
||||
A financial research agent that demonstrates structured research workflows with agents and tools for financial data analysis.
|
||||
|
||||
- **[handoffs](https://github.com/openai/openai-agents-python/tree/main/examples/handoffs):**
|
||||
See practical examples of agent handoffs.
|
||||
- **[handoffs](https://github.com/openai/openai-agents-python/tree/main/examples/handoffs):**
|
||||
Practical examples of agent handoffs with message filtering, including:
|
||||
|
||||
- **[mcp](https://github.com/openai/openai-agents-python/tree/main/examples/mcp):**
|
||||
Learn how to build agents with MCP.
|
||||
- Message filter example (`examples/handoffs/message_filter.py`)
|
||||
- Message filter with streaming (`examples/handoffs/message_filter_streaming.py`)
|
||||
|
||||
- **[customer_service](https://github.com/openai/openai-agents-python/tree/main/examples/customer_service)** and **[research_bot](https://github.com/openai/openai-agents-python/tree/main/examples/research_bot):**
|
||||
Two more built-out examples that illustrate real-world applications
|
||||
- **[hosted_mcp](https://github.com/openai/openai-agents-python/tree/main/examples/hosted_mcp):**
|
||||
Examples demonstrating how to use hosted MCP (Model Context Protocol) with the OpenAI Responses API, including:
|
||||
|
||||
- **customer_service**: Example customer service system for an airline.
|
||||
- **research_bot**: Simple deep research clone.
|
||||
- Simple hosted MCP without approval (`examples/hosted_mcp/simple.py`)
|
||||
- MCP connectors such as Google Calendar (`examples/hosted_mcp/connectors.py`)
|
||||
- Human-in-the-loop with interruption-based approvals (`examples/hosted_mcp/human_in_the_loop.py`)
|
||||
- On-approval callback for MCP tool calls (`examples/hosted_mcp/on_approval.py`)
|
||||
|
||||
- **[voice](https://github.com/openai/openai-agents-python/tree/main/examples/voice):**
|
||||
See examples of voice agents, using our TTS and STT models.
|
||||
- **[mcp](https://github.com/openai/openai-agents-python/tree/main/examples/mcp):**
|
||||
Learn how to build agents with MCP (Model Context Protocol), including:
|
||||
|
||||
- Filesystem examples
|
||||
- Git examples
|
||||
- MCP prompt server examples
|
||||
- SSE (Server-Sent Events) examples
|
||||
- SSE remote server connection (`examples/mcp/sse_remote_example`)
|
||||
- Streamable HTTP examples
|
||||
- Streamable HTTP remote connection (`examples/mcp/streamable_http_remote_example`)
|
||||
- Custom HTTP client factory for Streamable HTTP (`examples/mcp/streamablehttp_custom_client_example`)
|
||||
- Prefetching all MCP tools with `MCPUtil.get_all_function_tools` (`examples/mcp/get_all_mcp_tools_example`)
|
||||
- MCPServerManager with FastAPI (`examples/mcp/manager_example`)
|
||||
- MCP tool filtering (`examples/mcp/tool_filter_example`)
|
||||
|
||||
- **[memory](https://github.com/openai/openai-agents-python/tree/main/examples/memory):**
|
||||
Examples of different memory implementations for agents, including:
|
||||
|
||||
- SQLite session storage
|
||||
- Advanced SQLite session storage
|
||||
- Redis session storage
|
||||
- SQLAlchemy session storage
|
||||
- Dapr state store session storage
|
||||
- Encrypted session storage
|
||||
- OpenAI Conversations session storage
|
||||
- Responses compaction session storage
|
||||
- Stateless Responses compaction with `ModelSettings(store=False)` (`examples/memory/compaction_session_stateless_example.py`)
|
||||
- File-backed session storage (`examples/memory/file_session.py`)
|
||||
- File-backed session with human-in-the-loop (`examples/memory/file_hitl_example.py`)
|
||||
- SQLite in-memory session with human-in-the-loop (`examples/memory/memory_session_hitl_example.py`)
|
||||
- OpenAI Conversations session with human-in-the-loop (`examples/memory/openai_session_hitl_example.py`)
|
||||
- HITL approval/rejection scenario across sessions (`examples/memory/hitl_session_scenario.py`)
|
||||
|
||||
- **[model_providers](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers):**
|
||||
Explore how to use non-OpenAI models with the SDK, including custom providers and third-party adapters.
|
||||
|
||||
- **[realtime](https://github.com/openai/openai-agents-python/tree/main/examples/realtime):**
|
||||
Examples showing how to build real-time experiences using the SDK, including:
|
||||
|
||||
- Web application patterns with structured text and image messages
|
||||
- Command-line audio loops and playback handling
|
||||
- Twilio Media Streams integration over WebSocket
|
||||
- Twilio SIP integration using Realtime Calls API attach flows
|
||||
|
||||
- **[reasoning_content](https://github.com/openai/openai-agents-python/tree/main/examples/reasoning_content):**
|
||||
Examples demonstrating how to work with reasoning content, including:
|
||||
|
||||
- Reasoning content with the Runner API, streaming and non-streaming (`examples/reasoning_content/runner_example.py`)
|
||||
- Reasoning content with OSS models via OpenRouter (`examples/reasoning_content/gpt_oss_stream.py`)
|
||||
- Basic reasoning content example (`examples/reasoning_content/main.py`)
|
||||
|
||||
- **[research_bot](https://github.com/openai/openai-agents-python/tree/main/examples/research_bot):**
|
||||
Simple deep research clone that demonstrates complex multi-agent research workflows.
|
||||
|
||||
- **[tools](https://github.com/openai/openai-agents-python/tree/main/examples/tools):**
|
||||
Learn how to implement OAI hosted tools and experimental Codex tooling such as:
|
||||
|
||||
- Web search and web search with filters
|
||||
- File search
|
||||
- Code interpreter
|
||||
- Apply patch tool with file editing and approval (`examples/tools/apply_patch.py`)
|
||||
- Shell tool execution with approval callbacks (`examples/tools/shell.py`)
|
||||
- Shell tool with human-in-the-loop interruption-based approvals (`examples/tools/shell_human_in_the_loop.py`)
|
||||
- Hosted container shell with inline skills (`examples/tools/container_shell_inline_skill.py`)
|
||||
- Hosted container shell with skill references (`examples/tools/container_shell_skill_reference.py`)
|
||||
- Local shell with local skills (`examples/tools/local_shell_skill.py`)
|
||||
- Tool search with namespaces and deferred tools (`examples/tools/tool_search.py`)
|
||||
- Computer use
|
||||
- Image generation
|
||||
- Experimental Codex tool workflows (`examples/tools/codex.py`)
|
||||
- Experimental Codex same-thread workflows (`examples/tools/codex_same_thread.py`)
|
||||
|
||||
- **[voice](https://github.com/openai/openai-agents-python/tree/main/examples/voice):**
|
||||
See examples of voice agents, using our TTS and STT models, including streamed voice examples.
|
||||
|
||||
+77
-2
@@ -1,12 +1,22 @@
|
||||
# Guardrails
|
||||
|
||||
Guardrails run _in parallel_ to your agents, enabling you to do checks and validations of user input. For example, imagine you have an agent that uses a very smart (and hence slow/expensive) model to help with customer requests. You wouldn't want malicious users to ask the model to help them with their math homework. So, you can run a guardrail with a fast/cheap model. If the guardrail detects malicious usage, it can immediately raise an error, which stops the expensive model from running and saves you time/money.
|
||||
Guardrails enable you to do checks and validations of user input and agent output. For example, imagine you have an agent that uses a very smart (and hence slow/expensive) model to help with customer requests. You wouldn't want malicious users to ask the model to help them with their math homework. So, you can run a guardrail with a fast/cheap model. If the guardrail detects malicious usage, it can immediately raise an error and prevent the expensive model from running, saving you time and money (**when using blocking guardrails; for parallel guardrails, the expensive model may have already started running before the guardrail completes. See "Execution modes" below for details**).
|
||||
|
||||
There are two kinds of guardrails:
|
||||
|
||||
1. Input guardrails run on the initial user input
|
||||
2. Output guardrails run on the final agent output
|
||||
|
||||
## Workflow boundaries
|
||||
|
||||
Guardrails are attached to agents and tools, but they do not all run at the same points in a workflow:
|
||||
|
||||
- **Input guardrails** run only for the first agent in the chain.
|
||||
- **Output guardrails** run only for the agent that produces the final output.
|
||||
- **Tool guardrails** run on every custom function-tool invocation, with input guardrails before execution and output guardrails after execution.
|
||||
|
||||
If you need checks around each custom function-tool call in a workflow that includes managers, handoffs, or delegated specialists, use tool guardrails instead of relying only on agent-level input/output guardrails.
|
||||
|
||||
## Input guardrails
|
||||
|
||||
Input guardrails run in 3 steps:
|
||||
@@ -19,11 +29,19 @@ Input guardrails run in 3 steps:
|
||||
|
||||
Input guardrails are intended to run on user input, so an agent's guardrails only run if the agent is the *first* agent. You might wonder, why is the `guardrails` property on the agent instead of passed to `Runner.run`? It's because guardrails tend to be related to the actual Agent - you'd run different guardrails for different agents, so colocating the code is useful for readability.
|
||||
|
||||
### Execution modes
|
||||
|
||||
Input guardrails support two execution modes:
|
||||
|
||||
- **Parallel execution** (default, `run_in_parallel=True`): The guardrail runs concurrently with the agent's execution. This provides the best latency since both start at the same time. However, if the guardrail fails, the agent may have already consumed tokens and executed tools before being cancelled.
|
||||
|
||||
- **Blocking execution** (`run_in_parallel=False`): The guardrail runs and completes *before* the agent starts. If the guardrail tripwire is triggered, the agent never executes, preventing token consumption and tool execution. This is ideal for cost optimization and when you want to avoid potential side effects from tool calls.
|
||||
|
||||
## Output guardrails
|
||||
|
||||
Output guardrails run in 3 steps:
|
||||
|
||||
1. First, the guardrail receives the same input passed to the agent.
|
||||
1. First, the guardrail receives the output produced by the agent.
|
||||
2. Next, the guardrail function runs to produce a [`GuardrailFunctionOutput`][agents.guardrail.GuardrailFunctionOutput], which is then wrapped in an [`OutputGuardrailResult`][agents.guardrail.OutputGuardrailResult]
|
||||
3. Finally, we check if [`.tripwire_triggered`][agents.guardrail.GuardrailFunctionOutput.tripwire_triggered] is true. If true, an [`OutputGuardrailTripwireTriggered`][agents.exceptions.OutputGuardrailTripwireTriggered] exception is raised, so you can appropriately respond to the user or handle the exception.
|
||||
|
||||
@@ -31,6 +49,18 @@ Output guardrails run in 3 steps:
|
||||
|
||||
Output guardrails are intended to run on the final agent output, so an agent's guardrails only run if the agent is the *last* agent. Similar to the input guardrails, we do this because guardrails tend to be related to the actual Agent - you'd run different guardrails for different agents, so colocating the code is useful for readability.
|
||||
|
||||
Output guardrails always run after the agent completes, so they don't support the `run_in_parallel` parameter.
|
||||
|
||||
## Tool guardrails
|
||||
|
||||
Tool guardrails wrap **function tools** and let you validate or block tool calls before and after execution. They are configured on the tool itself and run every time that tool is invoked.
|
||||
|
||||
- Input tool guardrails run before the tool executes and can skip the call, replace the output with a message, or raise a tripwire.
|
||||
- Output tool guardrails run after the tool executes and can replace the output or raise a tripwire.
|
||||
- Tool guardrails apply only to function tools created with [`function_tool`][agents.tool.function_tool]. Handoffs run through the SDK's handoff pipeline rather than the normal function-tool pipeline, so tool guardrails do not apply to the handoff call itself. Hosted tools (`WebSearchTool`, `FileSearchTool`, `HostedMCPTool`, `CodeInterpreterTool`, `ImageGenerationTool`) and built-in execution tools (`ComputerTool`, `ShellTool`, `ApplyPatchTool`, `LocalShellTool`) also do not use this guardrail pipeline, and [`Agent.as_tool()`][agents.agent.Agent.as_tool] does not currently expose tool-guardrail options directly.
|
||||
|
||||
See the code snippet below for details.
|
||||
|
||||
## Tripwires
|
||||
|
||||
If the input or output fails the guardrail, the Guardrail can signal this with a tripwire. As soon as we see a guardrail that has triggered the tripwires, we immediately raise a `{Input,Output}GuardrailTripwireTriggered` exception and halt the Agent execution.
|
||||
@@ -152,3 +182,48 @@ async def main():
|
||||
2. This is the guardrail's output type.
|
||||
3. This is the guardrail function that receives the agent's output, and returns the result.
|
||||
4. This is the actual agent that defines the workflow.
|
||||
|
||||
Lastly, here are examples of tool guardrails.
|
||||
|
||||
```python
|
||||
import json
|
||||
from agents import (
|
||||
Agent,
|
||||
Runner,
|
||||
ToolGuardrailFunctionOutput,
|
||||
function_tool,
|
||||
tool_input_guardrail,
|
||||
tool_output_guardrail,
|
||||
)
|
||||
|
||||
@tool_input_guardrail
|
||||
def block_secrets(data):
|
||||
args = json.loads(data.context.tool_arguments or "{}")
|
||||
if "sk-" in json.dumps(args):
|
||||
return ToolGuardrailFunctionOutput.reject_content(
|
||||
"Remove secrets before calling this tool."
|
||||
)
|
||||
return ToolGuardrailFunctionOutput.allow()
|
||||
|
||||
|
||||
@tool_output_guardrail
|
||||
def redact_output(data):
|
||||
text = str(data.output or "")
|
||||
if "sk-" in text:
|
||||
return ToolGuardrailFunctionOutput.reject_content("Output contained sensitive data.")
|
||||
return ToolGuardrailFunctionOutput.allow()
|
||||
|
||||
|
||||
@function_tool(
|
||||
tool_input_guardrails=[block_secrets],
|
||||
tool_output_guardrails=[redact_output],
|
||||
)
|
||||
def classify_text(text: str) -> str:
|
||||
"""Classify text for internal routing."""
|
||||
return f"length:{len(text)}"
|
||||
|
||||
|
||||
agent = Agent(name="Classifier", tools=[classify_text])
|
||||
result = Runner.run_sync(agent, "hello world")
|
||||
print(result.final_output)
|
||||
```
|
||||
|
||||
+41
-2
@@ -8,9 +8,11 @@ Handoffs are represented as tools to the LLM. So if there's a handoff to an agen
|
||||
|
||||
All agents have a [`handoffs`][agents.agent.Agent.handoffs] param, which can either take an `Agent` directly, or a `Handoff` object that customizes the Handoff.
|
||||
|
||||
If you pass plain `Agent` instances, their [`handoff_description`][agents.agent.Agent.handoff_description] (when set) is appended to the default tool description. Use it to hint when the model should pick that handoff without writing a full `handoff()` object.
|
||||
|
||||
You can create a handoff using the [`handoff()`][agents.handoffs.handoff] function provided by the Agents SDK. This function allows you to specify the agent to hand off to, along with optional overrides and input filters.
|
||||
|
||||
### Basic Usage
|
||||
### Basic usage
|
||||
|
||||
Here's how you can create a simple handoff:
|
||||
|
||||
@@ -34,8 +36,12 @@ The [`handoff()`][agents.handoffs.handoff] function lets you customize things.
|
||||
- `tool_name_override`: By default, the `Handoff.default_tool_name()` function is used, which resolves to `transfer_to_<agent_name>`. You can override this.
|
||||
- `tool_description_override`: Override the default tool description from `Handoff.default_tool_description()`
|
||||
- `on_handoff`: A callback function executed when the handoff is invoked. This is useful for things like kicking off some data fetching as soon as you know a handoff is being invoked. This function receives the agent context, and can optionally also receive LLM generated input. The input data is controlled by the `input_type` param.
|
||||
- `input_type`: The type of input expected by the handoff (optional).
|
||||
- `input_type`: The schema for the handoff tool-call arguments. When set, the parsed payload is passed to `on_handoff`.
|
||||
- `input_filter`: This lets you filter the input received by the next agent. See below for more.
|
||||
- `is_enabled`: Whether the handoff is enabled. This can be a boolean or a function that returns a boolean, allowing you to dynamically enable or disable the handoff at runtime.
|
||||
- `nest_handoff_history`: Optional per-call override for the RunConfig-level `nest_handoff_history` setting. If `None`, the value defined in the active run configuration is used instead.
|
||||
|
||||
The [`handoff()`][agents.handoffs.handoff] helper always transfers control to the specific `agent` you passed in. If you have multiple possible destinations, register one handoff per destination and let the model choose among them. Use a custom [`Handoff`][agents.handoffs.Handoff] only when your own handoff code must decide which agent to return at invocation time.
|
||||
|
||||
```python
|
||||
from agents import Agent, handoff, RunContextWrapper
|
||||
@@ -77,10 +83,43 @@ handoff_obj = handoff(
|
||||
)
|
||||
```
|
||||
|
||||
`input_type` describes the arguments for the handoff tool call itself. The SDK exposes that schema to the model as the handoff tool's `parameters`, validates the returned JSON locally, and passes the parsed value to `on_handoff`.
|
||||
|
||||
It does not replace the next agent's main input, and it does not choose a different destination. The [`handoff()`][agents.handoffs.handoff] helper still transfers to the specific agent you wrapped, and the receiving agent still sees the conversation history unless you change it with an [`input_filter`][agents.handoffs.Handoff.input_filter] or nested handoff history settings.
|
||||
|
||||
`input_type` is also separate from [`RunContextWrapper.context`][agents.run_context.RunContextWrapper.context]. Use `input_type` for metadata the model decides at handoff time, not for application state or dependencies you already have locally.
|
||||
|
||||
### When to use `input_type`
|
||||
|
||||
Use `input_type` when the handoff needs a small piece of model-generated metadata such as `reason`, `language`, `priority`, or `summary`. For example, a triage agent can hand off to a refund agent with `{ "reason": "duplicate_charge", "priority": "high" }`, and `on_handoff` can log or persist that metadata before the refund agent takes over.
|
||||
|
||||
Choose a different mechanism when the goal is different:
|
||||
|
||||
- Put existing application state and dependencies in [`RunContextWrapper.context`][agents.run_context.RunContextWrapper.context]. See the [context guide](context.md).
|
||||
- Use [`input_filter`][agents.handoffs.Handoff.input_filter], [`RunConfig.nest_handoff_history`][agents.run.RunConfig.nest_handoff_history], or [`RunConfig.handoff_history_mapper`][agents.run.RunConfig.handoff_history_mapper] if you want to change what history the receiving agent sees.
|
||||
- Register one handoff per destination if there are multiple possible specialists. `input_type` can add metadata to the chosen handoff, but it does not dispatch between destinations.
|
||||
- If you want structured input for a nested specialist without transferring the conversation, prefer [`Agent.as_tool(parameters=...)`][agents.agent.Agent.as_tool]. See [tools](tools.md#structured-input-for-tool-agents).
|
||||
|
||||
## Input filters
|
||||
|
||||
When a handoff occurs, it's as though the new agent takes over the conversation, and gets to see the entire previous conversation history. If you want to change this, you can set an [`input_filter`][agents.handoffs.Handoff.input_filter]. An input filter is a function that receives the existing input via a [`HandoffInputData`][agents.handoffs.HandoffInputData], and must return a new `HandoffInputData`.
|
||||
|
||||
[`HandoffInputData`][agents.handoffs.HandoffInputData] includes:
|
||||
|
||||
- `input_history`: the input history before `Runner.run(...)` started.
|
||||
- `pre_handoff_items`: items generated before the agent turn where the handoff was invoked.
|
||||
- `new_items`: items generated during the current turn, including the handoff call and handoff output items.
|
||||
- `input_items`: optional items to forward to the next agent instead of `new_items`, allowing you to filter model input while keeping `new_items` intact for session history.
|
||||
- `run_context`: the active [`RunContextWrapper`][agents.run_context.RunContextWrapper] at the time the handoff was invoked.
|
||||
|
||||
Nested handoffs are available as an opt-in beta and are disabled by default while we stabilize them. When you enable [`RunConfig.nest_handoff_history`][agents.run.RunConfig.nest_handoff_history], the runner collapses the prior transcript into a single assistant summary message and wraps it in a `<CONVERSATION HISTORY>` block that keeps appending new turns when multiple handoffs happen during the same run. You can provide your own mapping function via [`RunConfig.handoff_history_mapper`][agents.run.RunConfig.handoff_history_mapper] to replace the generated message without writing a full `input_filter`. The opt-in only applies when neither the handoff nor the run supplies an explicit `input_filter`, so existing code that already customizes the payload (including the examples in this repository) keeps its current behavior without changes. You can override the nesting behaviour for a single handoff by passing `nest_handoff_history=True` or `False` to [`handoff(...)`][agents.handoffs.handoff], which sets [`Handoff.nest_handoff_history`][agents.handoffs.Handoff.nest_handoff_history]. If you just need to change the wrapper text for the generated summary, call [`set_conversation_history_wrappers`][agents.handoffs.set_conversation_history_wrappers] (and optionally [`reset_conversation_history_wrappers`][agents.handoffs.reset_conversation_history_wrappers]) before running your agents.
|
||||
|
||||
If both the handoff and the active [`RunConfig.handoff_input_filter`][agents.run.RunConfig.handoff_input_filter] define a filter, the per-handoff [`input_filter`][agents.handoffs.Handoff.input_filter] takes precedence for that specific handoff.
|
||||
|
||||
!!! note
|
||||
|
||||
Handoffs stay within a single run. Input guardrails still apply only to the first agent in the chain, and output guardrails only to the agent that produces the final output. Use tool guardrails when you need checks around each custom function-tool call inside the workflow.
|
||||
|
||||
There are some common patterns (for example removing all tool calls from the history), which are implemented for you in [`agents.extensions.handoff_filters`][]
|
||||
|
||||
```python
|
||||
|
||||
@@ -0,0 +1,205 @@
|
||||
# Human-in-the-loop
|
||||
|
||||
Use the human-in-the-loop (HITL) flow to pause agent execution until a person approves or rejects sensitive tool calls. Tools declare when they need approval, run results surface pending approvals as interruptions, and `RunState` lets you serialize and resume runs after decisions are made.
|
||||
|
||||
That approval surface is run-wide, not limited to the current top-level agent. The same pattern applies when the tool belongs to the current agent, to an agent reached through a handoff, or to a nested [`Agent.as_tool()`][agents.agent.Agent.as_tool] execution. In the nested `Agent.as_tool()` case, the interruption still surfaces on the outer run, so you approve or reject it on the outer `RunState` and resume the original top-level run.
|
||||
|
||||
With `Agent.as_tool()`, approvals can happen at two different layers: the agent tool itself can require approval via `Agent.as_tool(..., needs_approval=...)`, and tools inside the nested agent can later raise their own approvals after the nested run starts. Both are handled through the same outer-run interruption flow.
|
||||
|
||||
This page focuses on the manual approval flow via `interruptions`. If your app can decide in code, some tool types also support programmatic approval callbacks so the run can continue without pausing.
|
||||
|
||||
## Marking tools that need approval
|
||||
|
||||
Set `needs_approval` to `True` to always require approval or provide an async function that decides per call. The callable receives the run context, parsed tool parameters, and the tool call ID.
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner, function_tool
|
||||
|
||||
|
||||
@function_tool(needs_approval=True)
|
||||
async def cancel_order(order_id: int) -> str:
|
||||
return f"Cancelled order {order_id}"
|
||||
|
||||
|
||||
async def requires_review(_ctx, params, _call_id) -> bool:
|
||||
return "refund" in params.get("subject", "").lower()
|
||||
|
||||
|
||||
@function_tool(needs_approval=requires_review)
|
||||
async def send_email(subject: str, body: str) -> str:
|
||||
return f"Sent '{subject}'"
|
||||
|
||||
|
||||
agent = Agent(
|
||||
name="Support agent",
|
||||
instructions="Handle tickets and ask for approval when needed.",
|
||||
tools=[cancel_order, send_email],
|
||||
)
|
||||
```
|
||||
|
||||
`needs_approval` is available on [`function_tool`][agents.tool.function_tool], [`Agent.as_tool`][agents.agent.Agent.as_tool], [`ShellTool`][agents.tool.ShellTool], and [`ApplyPatchTool`][agents.tool.ApplyPatchTool]. Local MCP servers also support approvals through `require_approval` on [`MCPServerStdio`][agents.mcp.server.MCPServerStdio], [`MCPServerSse`][agents.mcp.server.MCPServerSse], and [`MCPServerStreamableHttp`][agents.mcp.server.MCPServerStreamableHttp]. Hosted MCP servers support approvals via [`HostedMCPTool`][agents.tool.HostedMCPTool] with `tool_config={"require_approval": "always"}` and an optional `on_approval_request` callback. Shell and apply_patch tools accept an `on_approval` callback if you want to auto-approve or auto-reject without surfacing an interruption.
|
||||
|
||||
## How the approval flow works
|
||||
|
||||
1. When the model emits a tool call, the runner evaluates its approval rule (`needs_approval`, `require_approval`, or the hosted MCP equivalent).
|
||||
2. If an approval decision for that tool call is already stored in the [`RunContextWrapper`][agents.run_context.RunContextWrapper], the runner proceeds without prompting. Per-call approvals are scoped to the specific call ID; pass `always_approve=True` or `always_reject=True` to persist the same decision for future calls to that tool during the rest of the run.
|
||||
3. Otherwise, execution pauses and `RunResult.interruptions` (or `RunResultStreaming.interruptions`) contains [`ToolApprovalItem`][agents.items.ToolApprovalItem] entries with details such as `agent.name`, `tool_name`, and `arguments`. This includes approvals raised after a handoff or inside nested `Agent.as_tool()` executions.
|
||||
4. Convert the result to a `RunState` with `result.to_state()`, call `state.approve(...)` or `state.reject(...)`, and then resume with `Runner.run(agent, state)` or `Runner.run_streamed(agent, state)`, where `agent` is the original top-level agent for the run.
|
||||
5. The resumed run continues where it left off and will re-enter this flow if new approvals are needed.
|
||||
|
||||
Sticky decisions created with `always_approve=True` or `always_reject=True` are stored in the run state, so they survive `state.to_string()` / `RunState.from_string(...)` and `state.to_json()` / `RunState.from_json(...)` when you resume the same paused run later.
|
||||
|
||||
You do not need to resolve every pending approval in the same pass. `interruptions` can contain a mix of regular function tools, hosted MCP approvals, and nested `Agent.as_tool()` approvals. If you rerun after approving or rejecting only some items, those resolved calls can continue while unresolved ones remain in `interruptions` and pause the run again.
|
||||
|
||||
## Custom rejection messages
|
||||
|
||||
By default, a rejected tool call returns the SDK's standard rejection text back into the run. You can customize that message in two layers:
|
||||
|
||||
- Run-wide fallback: set [`RunConfig.tool_error_formatter`][agents.run.RunConfig.tool_error_formatter] to control the default model-visible message for approval rejections across the whole run.
|
||||
- Per-call override: pass `rejection_message=...` to `state.reject(...)` when you want one specific rejected tool call to surface a different message.
|
||||
|
||||
If both are provided, the per-call `rejection_message` takes precedence over the run-wide formatter.
|
||||
|
||||
```python
|
||||
from agents import RunConfig, ToolErrorFormatterArgs
|
||||
|
||||
|
||||
def format_rejection(args: ToolErrorFormatterArgs[None]) -> str | None:
|
||||
if args.kind != "approval_rejected":
|
||||
return None
|
||||
return "Publish action was canceled because approval was rejected."
|
||||
|
||||
|
||||
run_config = RunConfig(tool_error_formatter=format_rejection)
|
||||
|
||||
# Later, while resolving a specific interruption:
|
||||
state.reject(
|
||||
interruption,
|
||||
rejection_message="Publish action was canceled because the reviewer denied approval.",
|
||||
)
|
||||
```
|
||||
|
||||
See [`examples/agent_patterns/human_in_the_loop_custom_rejection.py`](https://github.com/openai/openai-agents-python/tree/main/examples/agent_patterns/human_in_the_loop_custom_rejection.py) for a complete example that shows both layers together.
|
||||
|
||||
## Automatic approval decisions
|
||||
|
||||
Manual `interruptions` are the most general pattern, but they are not the only one:
|
||||
|
||||
- Local [`ShellTool`][agents.tool.ShellTool] and [`ApplyPatchTool`][agents.tool.ApplyPatchTool] can use `on_approval` to approve or reject immediately in code.
|
||||
- [`HostedMCPTool`][agents.tool.HostedMCPTool] can use `tool_config={"require_approval": "always"}` together with `on_approval_request` for the same kind of programmatic decision.
|
||||
- Plain [`function_tool`][agents.tool.function_tool] tools and [`Agent.as_tool()`][agents.agent.Agent.as_tool] use the manual interruption flow on this page.
|
||||
|
||||
When these callbacks return a decision, the run continues without pausing for a human response. For Realtime and voice session APIs, see the approval flow in the [Realtime guide](realtime/guide.md).
|
||||
|
||||
## Streaming and sessions
|
||||
|
||||
The same interruption flow works in streaming runs. After a streamed run pauses, keep consuming [`RunResultStreaming.stream_events()`][agents.result.RunResultStreaming.stream_events] until the iterator finishes, inspect [`RunResultStreaming.interruptions`][agents.result.RunResultStreaming.interruptions], resolve them, and resume with [`Runner.run_streamed(...)`][agents.run.Runner.run_streamed] if you want the resumed output to keep streaming. See [Streaming](streaming.md) for the streamed version of this pattern.
|
||||
|
||||
If you are also using a session, keep passing the same session instance when you resume from `RunState`, or pass another session object that points at the same backing store. The resumed turn is then appended to the same stored conversation history. See [Sessions](sessions/index.md) for the session lifecycle details.
|
||||
|
||||
## Example: pause, approve, resume
|
||||
|
||||
The snippet below mirrors the JavaScript HITL guide: it pauses when a tool needs approval, persists state to disk, reloads it, and resumes after collecting a decision.
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
import json
|
||||
from pathlib import Path
|
||||
|
||||
from agents import Agent, Runner, RunState, function_tool
|
||||
|
||||
|
||||
async def needs_oakland_approval(_ctx, params, _call_id) -> bool:
|
||||
return "Oakland" in params.get("city", "")
|
||||
|
||||
|
||||
@function_tool(needs_approval=needs_oakland_approval)
|
||||
async def get_temperature(city: str) -> str:
|
||||
return f"The temperature in {city} is 20° Celsius"
|
||||
|
||||
|
||||
agent = Agent(
|
||||
name="Weather assistant",
|
||||
instructions="Answer weather questions with the provided tools.",
|
||||
tools=[get_temperature],
|
||||
)
|
||||
|
||||
STATE_PATH = Path(".cache/hitl_state.json")
|
||||
|
||||
|
||||
def prompt_approval(tool_name: str, arguments: str | None) -> bool:
|
||||
answer = input(f"Approve {tool_name} with {arguments}? [y/N]: ").strip().lower()
|
||||
return answer in {"y", "yes"}
|
||||
|
||||
|
||||
async def main() -> None:
|
||||
result = await Runner.run(agent, "What is the temperature in Oakland?")
|
||||
|
||||
while result.interruptions:
|
||||
# Persist the paused state.
|
||||
state = result.to_state()
|
||||
STATE_PATH.parent.mkdir(parents=True, exist_ok=True)
|
||||
STATE_PATH.write_text(state.to_string())
|
||||
|
||||
# Load the state later (could be a different process).
|
||||
stored = json.loads(STATE_PATH.read_text())
|
||||
state = await RunState.from_json(agent, stored)
|
||||
|
||||
for interruption in result.interruptions:
|
||||
approved = await asyncio.get_running_loop().run_in_executor(
|
||||
None, prompt_approval, interruption.name or "unknown_tool", interruption.arguments
|
||||
)
|
||||
if approved:
|
||||
state.approve(interruption, always_approve=False)
|
||||
else:
|
||||
state.reject(interruption)
|
||||
|
||||
result = await Runner.run(agent, state)
|
||||
|
||||
print(result.final_output)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
In this example, `prompt_approval` is synchronous because it uses `input()` and is executed with `run_in_executor(...)`. If your approval source is already asynchronous (for example, an HTTP request or async database query), you can use an `async def` function and `await` it directly instead.
|
||||
|
||||
To stream output while waiting for approvals, call `Runner.run_streamed`, consume `result.stream_events()` until it completes, and then follow the same `result.to_state()` and resume steps shown above.
|
||||
|
||||
## Repository patterns and examples
|
||||
|
||||
- **Streaming approvals**: `examples/agent_patterns/human_in_the_loop_stream.py` shows how to drain `stream_events()` and then approve pending tool calls before resuming with `Runner.run_streamed(agent, state)`.
|
||||
- **Custom rejection text**: `examples/agent_patterns/human_in_the_loop_custom_rejection.py` shows how to combine run-level `tool_error_formatter` with per-call `rejection_message` overrides when approvals are rejected.
|
||||
- **Agent as tool approvals**: `Agent.as_tool(..., needs_approval=...)` applies the same interruption flow when delegated agent tasks need review. Nested interruptions still surface on the outer run, so resume the original top-level agent rather than the nested one.
|
||||
- **Local shell and apply_patch tools**: `ShellTool` and `ApplyPatchTool` also support `needs_approval`. Use `state.approve(interruption, always_approve=True)` or `state.reject(..., always_reject=True)` to cache the decision for future calls. For automatic decisions, provide `on_approval` (see `examples/tools/shell.py`); for manual decisions, handle interruptions (see `examples/tools/shell_human_in_the_loop.py`). Hosted shell environments do not support `needs_approval` or `on_approval`; see the [tools guide](tools.md).
|
||||
- **Local MCP servers**: Use `require_approval` on `MCPServerStdio` / `MCPServerSse` / `MCPServerStreamableHttp` to gate MCP tool calls (see `examples/mcp/get_all_mcp_tools_example/main.py` and `examples/mcp/tool_filter_example/main.py`).
|
||||
- **Hosted MCP servers**: Set `require_approval` to `"always"` on `HostedMCPTool` to force HITL, optionally providing `on_approval_request` to auto-approve or reject (see `examples/hosted_mcp/human_in_the_loop.py` and `examples/hosted_mcp/on_approval.py`). Use `"never"` for trusted servers (`examples/hosted_mcp/simple.py`).
|
||||
- **Sessions and memory**: Pass a session to `Runner.run` so approvals and conversation history survive multiple turns. SQLite and OpenAI Conversations session variants are in `examples/memory/memory_session_hitl_example.py` and `examples/memory/openai_session_hitl_example.py`.
|
||||
- **Realtime agents**: The realtime demo exposes WebSocket messages that approve or reject tool calls via `approve_tool_call` / `reject_tool_call` on the `RealtimeSession` (see `examples/realtime/app/server.py` for the server-side handlers and [Realtime guide](realtime/guide.md#tool-approvals) for the API surface).
|
||||
|
||||
## Long-running approvals
|
||||
|
||||
`RunState` is designed to be durable. Use `state.to_json()` or `state.to_string()` to store pending work in a database or queue and recreate it later with `RunState.from_json(...)` or `RunState.from_string(...)`.
|
||||
|
||||
Useful serialization options:
|
||||
|
||||
- `context_serializer`: Customize how non-mapping context objects are serialized.
|
||||
- `context_deserializer`: Rebuild non-mapping context objects when loading state with `RunState.from_json(...)` or `RunState.from_string(...)`.
|
||||
- `strict_context=True`: Fail serialization or deserialization unless the context is already a
|
||||
mapping or you provide the appropriate serializer/deserializer.
|
||||
- `context_override`: Replace the serialized context when loading state. This is useful when you
|
||||
do not want to restore the original context object, but it does not remove that context from an
|
||||
already serialized payload.
|
||||
- `include_tracing_api_key=True`: Include the tracing API key in the serialized trace payload
|
||||
when you need resumed work to keep exporting traces with the same credentials.
|
||||
|
||||
Serialized run state includes your app context plus SDK-managed runtime metadata such as approvals,
|
||||
usage, serialized `tool_input`, nested agent-as-tool resumptions, trace metadata, and server-managed
|
||||
conversation settings. If you plan to store or transmit serialized state, treat
|
||||
`RunContextWrapper.context` as persisted data and avoid placing secrets there unless you
|
||||
intentionally want them to travel with the state.
|
||||
|
||||
## Versioning pending tasks
|
||||
|
||||
If approvals may sit for a while, store a version marker for your agent definitions or SDK alongside the serialized state. You can then route deserialization to the matching code path to avoid incompatibilities when models, prompts, or tool definitions change.
|
||||
+53
-8
@@ -3,8 +3,8 @@
|
||||
The [OpenAI Agents SDK](https://github.com/openai/openai-agents-python) enables you to build agentic AI apps in a lightweight, easy-to-use package with very few abstractions. It's a production-ready upgrade of our previous experimentation for agents, [Swarm](https://github.com/openai/swarm/tree/main). The Agents SDK has a very small set of primitives:
|
||||
|
||||
- **Agents**, which are LLMs equipped with instructions and tools
|
||||
- **Handoffs**, which allow agents to delegate to other agents for specific tasks
|
||||
- **Guardrails**, which enable the inputs to agents to be validated
|
||||
- **Agents as tools / Handoffs**, which allow agents to delegate to other agents for specific tasks
|
||||
- **Guardrails**, which enable validation of agent inputs and outputs
|
||||
|
||||
In combination with Python, these primitives are powerful enough to express complex relationships between tools and agents, and allow you to build real-world applications without a steep learning curve. In addition, the SDK comes with built-in **tracing** that lets you visualize and debug your agentic flows, as well as evaluate them and even fine-tune models for your application.
|
||||
|
||||
@@ -17,12 +17,34 @@ The SDK has two driving design principles:
|
||||
|
||||
Here are the main features of the SDK:
|
||||
|
||||
- Agent loop: Built-in agent loop that handles calling tools, sending results to the LLM, and looping until the LLM is done.
|
||||
- Python-first: Use built-in language features to orchestrate and chain agents, rather than needing to learn new abstractions.
|
||||
- Handoffs: A powerful feature to coordinate and delegate between multiple agents.
|
||||
- Guardrails: Run input validations and checks in parallel to your agents, breaking early if the checks fail.
|
||||
- Function tools: Turn any Python function into a tool, with automatic schema generation and Pydantic-powered validation.
|
||||
- Tracing: Built-in tracing that lets you visualize, debug and monitor your workflows, as well as use the OpenAI suite of evaluation, fine-tuning and distillation tools.
|
||||
- **Agent loop**: A built-in agent loop that handles tool invocation, sends results back to the LLM, and continues until the task is complete.
|
||||
- **Python-first**: Use built-in language features to orchestrate and chain agents, rather than needing to learn new abstractions.
|
||||
- **Agents as tools / Handoffs**: A powerful mechanism for coordinating and delegating work across multiple agents.
|
||||
- **Sandbox agents**: Run specialists inside real isolated workspaces with manifest-defined files, sandbox client choice, and resumable sandbox sessions.
|
||||
- **Guardrails**: Run input validation and safety checks in parallel with agent execution, and fail fast when checks do not pass.
|
||||
- **Function tools**: Turn any Python function into a tool with automatic schema generation and Pydantic-powered validation.
|
||||
- **MCP server tool calling**: Built-in MCP server tool integration that works the same way as function tools.
|
||||
- **Sessions**: A persistent memory layer for maintaining working context within an agent loop.
|
||||
- **Human in the loop**: Built-in mechanisms for involving humans across agent runs.
|
||||
- **Tracing**: Built-in tracing for visualizing, debugging, and monitoring workflows, with support for the OpenAI suite of evaluation, fine-tuning, and distillation tools.
|
||||
- **Realtime Agents**: Build powerful voice agents with `gpt-realtime-2`, automatic interruption detection, context management, guardrails, and more.
|
||||
|
||||
## Agents SDK or Responses API?
|
||||
|
||||
The SDK uses the Responses API by default for OpenAI models, but it adds a higher-level runtime around model calls.
|
||||
|
||||
Use the Responses API directly when:
|
||||
|
||||
- you want to own the loop, tool dispatch, and state handling yourself
|
||||
- your workflow is short-lived and mainly about returning the model's response
|
||||
|
||||
Use the Agents SDK when:
|
||||
|
||||
- you want the runtime to manage turns, tool execution, guardrails, handoffs, or sessions
|
||||
- your agent should produce artifacts or operate across multiple coordinated steps
|
||||
- you need a real workspace or resumable execution through [Sandbox agents](sandbox_agents.md)
|
||||
|
||||
You do not need to choose one globally. Many applications use the SDK for managed workflows and call the Responses API directly for lower-level paths.
|
||||
|
||||
## Installation
|
||||
|
||||
@@ -50,3 +72,26 @@ print(result.final_output)
|
||||
```bash
|
||||
export OPENAI_API_KEY=sk-...
|
||||
```
|
||||
|
||||
## Start here
|
||||
|
||||
- Build your first text-based agent with the [Quickstart](quickstart.md).
|
||||
- Then decide how you want to carry state across turns in [Running agents](running_agents.md#choose-a-memory-strategy).
|
||||
- If the task depends on real files, repos, or isolated per-agent workspace state, read the [Sandbox agents quickstart](sandbox_agents.md).
|
||||
- If you are deciding between handoffs and manager-style orchestration, read [Agent orchestration](multi_agent.md).
|
||||
|
||||
## Choose your path
|
||||
|
||||
Use this table when you know the job you want to do, but not which page explains it.
|
||||
|
||||
| Goal | Start here |
|
||||
| --- | --- |
|
||||
| Build the first text agent and see one complete run | [Quickstart](quickstart.md) |
|
||||
| Add function tools, hosted tools, or agents as tools | [Tools](tools.md) |
|
||||
| Run a coding, review, or document agent inside a real isolated workspace | [Sandbox agents quickstart](sandbox_agents.md) and [Sandbox clients](sandbox/clients.md) |
|
||||
| Decide between handoffs and manager-style orchestration | [Agent orchestration](multi_agent.md) |
|
||||
| Keep memory across turns | [Running agents](running_agents.md#choose-a-memory-strategy) and [Sessions](sessions/index.md) |
|
||||
| Use OpenAI models, websocket transport, or non-OpenAI providers | [Models](models/index.md) |
|
||||
| Review outputs, run items, interruptions, and resume state | [Results](results.md) |
|
||||
| Build a low-latency voice agent with `gpt-realtime-2` | [Realtime agents quickstart](realtime/quickstart.md) and [Realtime transport](realtime/transport.md) |
|
||||
| Build a speech-to-text / agent / text-to-speech pipeline | [Voice pipeline quickstart](voice/quickstart.md) |
|
||||
|
||||
+312
-30
@@ -1,37 +1,140 @@
|
||||
---
|
||||
search:
|
||||
exclude: true
|
||||
---
|
||||
# エージェント
|
||||
|
||||
エージェントは、アプリケーションの中核となる基本コンポーネントです。エージェントとは、instructions とツールで構成された大規模言語モデル(LLM)のことです。
|
||||
エージェントは、アプリにおける中核的な構成要素です。エージェントとは、instructions、ツール、およびハンドオフ、ガードレール、structured outputs などの任意のランタイム動作で設定された大規模言語モデル (LLM) です。
|
||||
|
||||
単一のシンプルな `Agent` を定義またはカスタマイズしたい場合は、このページを使用してください。複数のエージェントをどのように連携させるかを検討している場合は、[エージェントオーケストレーション](multi_agent.md) を読んでください。エージェントを、マニフェストで定義されたファイルとサンドボックスネイティブ機能を備えた分離ワークスペース内で実行する必要がある場合は、[Sandbox エージェントの概念](sandbox/guide.md) を読んでください。
|
||||
|
||||
SDK は OpenAI モデルに対してデフォルトで Responses API を使用しますが、ここでの違いはオーケストレーションです。`Agent` と `Runner` により、SDK がターン、ツール、ガードレール、ハンドオフ、セッションを管理できます。そのループを自分で管理したい場合は、代わりに Responses API を直接使用してください。
|
||||
|
||||
## 次のガイドの選択
|
||||
|
||||
このページをエージェント定義のハブとして使用してください。次に行う判断に合ったガイドへ進んでください。
|
||||
|
||||
| したいこと | 次に読むもの |
|
||||
| --- | --- |
|
||||
| モデルまたはプロバイダー設定を選択する | [モデル](models/index.md) |
|
||||
| エージェントに機能を追加する | [ツール](tools.md) |
|
||||
| 実際のリポジトリ、ドキュメントバンドル、または分離ワークスペースに対してエージェントを実行する | [Sandbox エージェントのクイックスタート](sandbox_agents.md) |
|
||||
| マネージャースタイルのオーケストレーションとハンドオフのどちらにするかを決める | [エージェントオーケストレーション](multi_agent.md) |
|
||||
| ハンドオフ動作を設定する | [ハンドオフ](handoffs.md) |
|
||||
| ターンを実行し、イベントをストリーミングし、会話状態を管理する | [エージェント実行](running_agents.md) |
|
||||
| 最終出力、実行アイテム、または再開可能な状態を確認する | [結果](results.md) |
|
||||
| ローカル依存関係とランタイム状態を共有する | [コンテキスト管理](context.md) |
|
||||
|
||||
## 基本設定
|
||||
|
||||
エージェントで最も一般的に設定するプロパティは以下の通りです。
|
||||
エージェントで最も一般的なプロパティは次のとおりです。
|
||||
|
||||
- `instructions`:developer message や システムプロンプト(system prompt)とも呼ばれます。
|
||||
- `model`:どの LLM を使用するか、また `model_settings` で temperature や top_p などのモデル調整パラメーターを設定できます。
|
||||
- `tools`:エージェントがタスクを達成するために使用できるツールです。
|
||||
| プロパティ | 必須 | 説明 |
|
||||
| --- | --- | --- |
|
||||
| `name` | はい | 人間が読みやすいエージェント名。 |
|
||||
| `instructions` | いいえ | システムプロンプトまたは動的 instructions コールバック。強く推奨します。[動的 instructions](#dynamic-instructions) を参照してください。 |
|
||||
| `prompt` | いいえ | OpenAI Responses API のプロンプト設定。静的なプロンプトオブジェクトまたは関数を受け付けます。[プロンプトテンプレート](#prompt-templates) を参照してください。 |
|
||||
| `handoff_description` | いいえ | このエージェントがハンドオフ先として提示されるときに公開される短い説明。 |
|
||||
| `handoffs` | いいえ | 会話を専門エージェントに委任します。[ハンドオフ](handoffs.md) を参照してください。 |
|
||||
| `model` | いいえ | 使用する LLM。[モデル](models/index.md) を参照してください。 |
|
||||
| `model_settings` | いいえ | `temperature`、`top_p`、`tool_choice` などのモデル調整パラメーター。 |
|
||||
| `tools` | いいえ | エージェントが呼び出せるツール。[ツール](tools.md) を参照してください。 |
|
||||
| `mcp_servers` | いいえ | エージェント向けの MCP ベースのツール。[MCP ガイド](mcp.md) を参照してください。 |
|
||||
| `mcp_config` | いいえ | 厳格なスキーマ変換や MCP 失敗時の形式設定など、MCP ツールの準備方法を微調整します。[MCP ガイド](mcp.md#agent-level-mcp-configuration) を参照してください。 |
|
||||
| `input_guardrails` | いいえ | このエージェントチェーンの最初のユーザー入力に対して実行されるガードレール。[ガードレール](guardrails.md) を参照してください。 |
|
||||
| `output_guardrails` | いいえ | このエージェントの最終出力に対して実行されるガードレール。[ガードレール](guardrails.md) を参照してください。 |
|
||||
| `output_type` | いいえ | プレーンテキストの代わりとなる構造化出力型。[出力型](#output-types) を参照してください。 |
|
||||
| `hooks` | いいえ | エージェントスコープのライフサイクルコールバック。[ライフサイクルイベント (フック)](#lifecycle-events-hooks) を参照してください。 |
|
||||
| `tool_use_behavior` | いいえ | ツールの結果をモデルに戻してループさせるか、実行を終了するかを制御します。[ツール使用の動作](#tool-use-behavior) を参照してください。 |
|
||||
| `reset_tool_choice` | いいえ | ツール呼び出し後に `tool_choice` をリセットします (デフォルト: `True`)。これによりツール使用ループを避けます。[ツール使用の強制](#forcing-tool-use) を参照してください。 |
|
||||
|
||||
```python
|
||||
from agents import Agent, ModelSettings, function_tool
|
||||
|
||||
@function_tool
|
||||
def get_weather(city: str) -> str:
|
||||
"""returns weather info for the specified city."""
|
||||
return f"The weather in {city} is sunny"
|
||||
|
||||
agent = Agent(
|
||||
name="Haiku agent",
|
||||
instructions="Always respond in haiku form",
|
||||
model="o3-mini",
|
||||
model="gpt-5-nano",
|
||||
tools=[get_weather],
|
||||
)
|
||||
```
|
||||
|
||||
このセクションの内容はすべて `Agent` に適用されます。`SandboxAgent` は同じ考え方を基盤とし、それに `default_manifest`、`base_instructions`、`capabilities`、`run_as` を加えて、ワークスペーススコープの実行に対応します。[Sandbox エージェントの概念](sandbox/guide.md) を参照してください。
|
||||
|
||||
## プロンプトテンプレート
|
||||
|
||||
`prompt` を設定することで、OpenAI プラットフォームで作成したプロンプトテンプレートを参照できます。これは Responses API を使用する OpenAI モデルで動作します。
|
||||
|
||||
使用するには、次の手順を行ってください。
|
||||
|
||||
1. https://platform.openai.com/playground/prompts に移動します
|
||||
2. 新しいプロンプト変数 `poem_style` を作成します。
|
||||
3. 次の内容でシステムプロンプトを作成します。
|
||||
|
||||
```
|
||||
Write a poem in {{poem_style}}
|
||||
```
|
||||
|
||||
4. この例を `--prompt-id` フラグ付きで実行します。
|
||||
|
||||
```python
|
||||
from agents import Agent
|
||||
|
||||
agent = Agent(
|
||||
name="Prompted assistant",
|
||||
prompt={
|
||||
"id": "pmpt_123",
|
||||
"version": "1",
|
||||
"variables": {"poem_style": "haiku"},
|
||||
},
|
||||
)
|
||||
```
|
||||
|
||||
実行時にプロンプトを動的に生成することもできます。
|
||||
|
||||
```python
|
||||
from dataclasses import dataclass
|
||||
|
||||
from agents import Agent, GenerateDynamicPromptData, Runner
|
||||
|
||||
@dataclass
|
||||
class PromptContext:
|
||||
prompt_id: str
|
||||
poem_style: str
|
||||
|
||||
|
||||
async def build_prompt(data: GenerateDynamicPromptData):
|
||||
ctx: PromptContext = data.context.context
|
||||
return {
|
||||
"id": ctx.prompt_id,
|
||||
"version": "1",
|
||||
"variables": {"poem_style": ctx.poem_style},
|
||||
}
|
||||
|
||||
|
||||
agent = Agent(name="Prompted assistant", prompt=build_prompt)
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
"Say hello",
|
||||
context=PromptContext(prompt_id="pmpt_123", poem_style="limerick"),
|
||||
)
|
||||
```
|
||||
|
||||
## コンテキスト
|
||||
|
||||
エージェントは `context` 型に対して汎用的です。コンテキストは依存性注入ツールであり、`Runner.run()` に渡すオブジェクトです。これはすべてのエージェント、ツール、ハンドオフなどに渡され、エージェント実行時の依存関係や状態をまとめて管理します。任意の Python オブジェクトを context として指定できます。
|
||||
エージェントは `context` 型に対してジェネリックです。コンテキストは依存性注入のためのツールです。自分で作成して `Runner.run()` に渡すオブジェクトであり、すべてのエージェント、ツール、ハンドオフなどに渡され、エージェント実行の依存関係や状態をまとめて保持します。コンテキストには任意の Python オブジェクトを指定できます。
|
||||
|
||||
完全な `RunContextWrapper` インターフェイス、共有された使用量追跡、ネストした `tool_input`、およびシリアライズ上の注意点については、[コンテキストガイド](context.md) を読んでください。
|
||||
|
||||
```python
|
||||
@dataclass
|
||||
class UserContext:
|
||||
name: str
|
||||
uid: str
|
||||
is_pro_user: bool
|
||||
|
||||
@@ -43,9 +146,9 @@ agent = Agent[UserContext](
|
||||
)
|
||||
```
|
||||
|
||||
## 出力タイプ
|
||||
## 出力型
|
||||
|
||||
デフォルトでは、エージェントはプレーンテキスト(つまり `str`)出力を生成します。特定の型の出力をエージェントに生成させたい場合は、`output_type` パラメーターを使用できます。一般的な選択肢として [Pydantic](https://docs.pydantic.dev/) オブジェクトがありますが、Pydantic の [TypeAdapter](https://docs.pydantic.dev/latest/api/type_adapter/) でラップできる型(dataclasses、リスト、TypedDict など)であればサポートしています。
|
||||
デフォルトでは、エージェントはプレーンテキスト (すなわち `str`) の出力を生成します。エージェントに特定の型の出力を生成させたい場合は、`output_type` パラメーターを使用できます。一般的な選択肢は [Pydantic](https://docs.pydantic.dev/) オブジェクトを使用することですが、Pydantic [TypeAdapter](https://docs.pydantic.dev/latest/api/type_adapter/) でラップできる任意の型にも対応しています。dataclasses、lists、TypedDict などです。
|
||||
|
||||
```python
|
||||
from pydantic import BaseModel
|
||||
@@ -66,11 +169,49 @@ agent = Agent(
|
||||
|
||||
!!! note
|
||||
|
||||
`output_type` を指定すると、モデルは通常のプレーンテキスト応答の代わりに [structured outputs](https://platform.openai.com/docs/guides/structured-outputs) を使用するよう指示されます。
|
||||
`output_type` を渡すと、モデルに通常のプレーンテキスト応答ではなく [structured outputs](https://platform.openai.com/docs/guides/structured-outputs) を使用するよう指示することになります。
|
||||
|
||||
## ハンドオフ
|
||||
## マルチエージェントシステムの設計パターン
|
||||
|
||||
ハンドオフは、エージェントが委任できるサブエージェントです。ハンドオフのリストを指定すると、エージェントは必要に応じてそれらに処理を委任できます。これは、単一タスクに特化したモジュール型のエージェントをオーケストレーションする強力なパターンです。詳細は [handoffs](handoffs.md) ドキュメントをご覧ください。
|
||||
マルチエージェントシステムを設計する方法は数多くありますが、幅広く適用できるパターンとして、よく見られるものが 2 つあります。
|
||||
|
||||
1. マネージャー (agents as tools): 中央のマネージャー/オーケストレーターが、専門化されたサブエージェントをツールとして呼び出し、会話の制御を維持します。
|
||||
2. ハンドオフ: 対等なエージェントが、会話を引き継ぐ専門エージェントへ制御をハンドオフします。これは分散型です。
|
||||
|
||||
詳細は、[エージェント構築の実践ガイド](https://cdn.openai.com/business-guides-and-resources/a-practical-guide-to-building-agents.pdf) を参照してください。
|
||||
|
||||
### マネージャー (agents as tools)
|
||||
|
||||
`customer_facing_agent` はユーザーとのやり取りをすべて処理し、ツールとして公開された専門サブエージェントを呼び出します。詳しくは [ツール](tools.md#agents-as-tools) ドキュメントを読んでください。
|
||||
|
||||
```python
|
||||
from agents import Agent
|
||||
|
||||
booking_agent = Agent(...)
|
||||
refund_agent = Agent(...)
|
||||
|
||||
customer_facing_agent = Agent(
|
||||
name="Customer-facing agent",
|
||||
instructions=(
|
||||
"Handle all direct user communication. "
|
||||
"Call the relevant tools when specialized expertise is needed."
|
||||
),
|
||||
tools=[
|
||||
booking_agent.as_tool(
|
||||
tool_name="booking_expert",
|
||||
tool_description="Handles booking questions and requests.",
|
||||
),
|
||||
refund_agent.as_tool(
|
||||
tool_name="refund_expert",
|
||||
tool_description="Handles refund questions and requests.",
|
||||
)
|
||||
],
|
||||
)
|
||||
```
|
||||
|
||||
### ハンドオフ
|
||||
|
||||
ハンドオフは、エージェントが委任できるサブエージェントです。ハンドオフが発生すると、委任先のエージェントが会話履歴を受け取り、会話を引き継ぎます。このパターンにより、単一タスクに優れた、モジュール化された専門エージェントを実現できます。詳しくは [ハンドオフ](handoffs.md) ドキュメントを読んでください。
|
||||
|
||||
```python
|
||||
from agents import Agent
|
||||
@@ -81,9 +222,9 @@ refund_agent = Agent(...)
|
||||
triage_agent = Agent(
|
||||
name="Triage agent",
|
||||
instructions=(
|
||||
"Help the user with their questions."
|
||||
"If they ask about booking, handoff to the booking agent."
|
||||
"If they ask about refunds, handoff to the refund agent."
|
||||
"Help the user with their questions. "
|
||||
"If they ask about booking, hand off to the booking agent. "
|
||||
"If they ask about refunds, hand off to the refund agent."
|
||||
),
|
||||
handoffs=[booking_agent, refund_agent],
|
||||
)
|
||||
@@ -91,7 +232,7 @@ triage_agent = Agent(
|
||||
|
||||
## 動的 instructions
|
||||
|
||||
多くの場合、エージェント作成時に instructions を指定できますが、関数を使って動的に instructions を提供することも可能です。この関数はエージェントと context を受け取り、プロンプトを返す必要があります。通常の関数と `async` 関数の両方が利用可能です。
|
||||
ほとんどの場合、エージェントを作成するときに instructions を指定できます。ただし、関数を通じて動的な instructions を提供することもできます。この関数はエージェントとコンテキストを受け取り、プロンプトを返す必要があります。通常の関数と `async` 関数のどちらも受け付けます。
|
||||
|
||||
```python
|
||||
def dynamic_instructions(
|
||||
@@ -106,23 +247,65 @@ agent = Agent[UserContext](
|
||||
)
|
||||
```
|
||||
|
||||
## ライフサイクルイベント(フック)
|
||||
## ライフサイクルイベント (フック)
|
||||
|
||||
エージェントのライフサイクルを監視したい場合があります。たとえば、イベントを記録したり、特定のイベント発生時にデータを事前取得したりしたい場合です。`hooks` プロパティを使ってエージェントのライフサイクルにフックできます。[`AgentHooks`][agents.lifecycle.AgentHooks] クラスをサブクラス化し、関心のあるメソッドをオーバーライドしてください。
|
||||
場合によっては、エージェントのライフサイクルを監視したいことがあります。たとえば、特定のイベントが発生したときに、イベントをログに記録したり、データを事前取得したり、使用量を記録したりできます。
|
||||
|
||||
フックのスコープは 2 つあります。
|
||||
|
||||
- [`RunHooks`][agents.lifecycle.RunHooks] は、他のエージェントへのハンドオフを含む `Runner.run(...)` 呼び出し全体を監視します。
|
||||
- [`AgentHooks`][agents.lifecycle.AgentHooks] は、`agent.hooks` を通じて特定のエージェントインスタンスにアタッチされます。
|
||||
|
||||
コールバックのコンテキストもイベントに応じて変わります。
|
||||
|
||||
- エージェントの開始/終了フックは [`AgentHookContext`][agents.run_context.AgentHookContext] を受け取ります。これは元のコンテキストをラップし、共有された実行使用量状態を保持します。
|
||||
- LLM、ツール、ハンドオフのフックは [`RunContextWrapper`][agents.run_context.RunContextWrapper] を受け取ります。
|
||||
|
||||
典型的なフックのタイミングは次のとおりです。
|
||||
|
||||
- `on_agent_start` / `on_agent_end`: 特定のエージェントが最終出力の生成を開始または終了するとき。
|
||||
- `on_llm_start` / `on_llm_end`: 各モデル呼び出しの直前/直後。
|
||||
- `on_tool_start` / `on_tool_end`: 各ローカルツール呼び出しの前後。
|
||||
関数ツールの場合、フックの `context` は通常 `ToolContext` なので、`tool_call_id` などのツール呼び出しメタデータを確認できます。
|
||||
- `on_handoff`: 制御があるエージェントから別のエージェントへ移るとき。
|
||||
|
||||
ワークフロー全体に対して単一のオブザーバーが必要な場合は `RunHooks` を使用し、1 つのエージェントにカスタム副作用が必要な場合は `AgentHooks` を使用します。
|
||||
|
||||
```python
|
||||
from agents import Agent, RunHooks, Runner
|
||||
|
||||
|
||||
class LoggingHooks(RunHooks):
|
||||
async def on_agent_start(self, context, agent):
|
||||
print(f"Starting {agent.name}")
|
||||
|
||||
async def on_llm_end(self, context, agent, response):
|
||||
print(f"{agent.name} produced {len(response.output)} output items")
|
||||
|
||||
async def on_agent_end(self, context, agent, output):
|
||||
print(f"{agent.name} finished with usage: {context.usage}")
|
||||
|
||||
|
||||
agent = Agent(name="Assistant", instructions="Be concise.")
|
||||
result = await Runner.run(agent, "Explain quines", hooks=LoggingHooks())
|
||||
print(result.final_output)
|
||||
```
|
||||
|
||||
完全なコールバック API については、[ライフサイクル API リファレンス](ref/lifecycle.md) を参照してください。
|
||||
|
||||
## ガードレール
|
||||
|
||||
ガードレールを使うと、エージェントの実行と並行して user 入力のチェックやバリデーションを行えます。たとえば、user の入力が関連性のある内容かどうかをスクリーニングできます。詳細は [guardrails](guardrails.md) ドキュメントをご覧ください。
|
||||
ガードレールを使うと、エージェントの実行と並行してユーザー入力に対するチェック/検証を実行し、エージェントの出力が生成された後にその出力に対するチェック/検証を実行できます。たとえば、ユーザー入力とエージェント出力について関連性をスクリーニングできます。詳しくは [ガードレール](guardrails.md) ドキュメントを読んでください。
|
||||
|
||||
## エージェントのクローン/コピー
|
||||
## エージェントのクローン/コピー
|
||||
|
||||
エージェントの `clone()` メソッドを使うことで、エージェントを複製し、任意のプロパティを変更できます。
|
||||
エージェントの `clone()` メソッドを使用すると、エージェントを複製し、必要に応じて任意のプロパティを変更できます。
|
||||
|
||||
```python
|
||||
pirate_agent = Agent(
|
||||
name="Pirate",
|
||||
instructions="Write like a pirate",
|
||||
model="o3-mini",
|
||||
model="gpt-5.5",
|
||||
)
|
||||
|
||||
robot_agent = pirate_agent.clone(
|
||||
@@ -133,15 +316,114 @@ robot_agent = pirate_agent.clone(
|
||||
|
||||
## ツール使用の強制
|
||||
|
||||
ツールのリストを指定しても、必ずしも LLM がツールを使用するとは限りません。[`ModelSettings.tool_choice`][agents.model_settings.ModelSettings.tool_choice] を設定することでツールの使用を強制できます。有効な値は以下の通りです。
|
||||
ツールのリストを指定しても、LLM が必ずツールを使用するとは限りません。[`ModelSettings.tool_choice`][agents.model_settings.ModelSettings.tool_choice] を設定することで、ツール使用を強制できます。有効な値は次のとおりです。
|
||||
|
||||
1. `auto`:LLM がツールを使うかどうかを自動で判断します。
|
||||
2. `required`:LLM にツールの使用を必須とします(どのツールを使うかは賢く選択されます)。
|
||||
3. `none`:LLM にツールを _使わない_ ことを要求します。
|
||||
4. 特定の文字列(例:`my_tool`)を指定すると、その特定のツールの使用を必須とします。
|
||||
1. `auto`: LLM がツールを使用するかどうかを判断できるようにします。
|
||||
2. `required`: LLM にツールの使用を要求します (ただし、どのツールを使うかは賢く判断できます)。
|
||||
3. `none`: LLM にツールを使用 _しない_ ことを要求します。
|
||||
4. `my_tool` などの特定の文字列を設定すると、LLM にその特定のツールを使用することを要求します。
|
||||
|
||||
OpenAI Responses のツール検索を使用している場合、名前付きツール選択にはより多くの制限があります。`tool_choice` で名前空間名だけや deferred-only ツールをターゲットにすることはできず、`tool_choice="tool_search"` は [`ToolSearchTool`][agents.tool.ToolSearchTool] をターゲットにしません。このような場合は、`auto` または `required` を優先してください。Responses 固有の制約については、[ホスト型ツール検索](tools.md#hosted-tool-search) を参照してください。
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner, function_tool, ModelSettings
|
||||
|
||||
@function_tool
|
||||
def get_weather(city: str) -> str:
|
||||
"""Returns weather info for the specified city."""
|
||||
return f"The weather in {city} is sunny"
|
||||
|
||||
agent = Agent(
|
||||
name="Weather Agent",
|
||||
instructions="Retrieve weather details.",
|
||||
tools=[get_weather],
|
||||
model_settings=ModelSettings(tool_choice="get_weather")
|
||||
)
|
||||
```
|
||||
|
||||
## ツール使用の動作
|
||||
|
||||
`Agent` 設定の `tool_use_behavior` パラメーターは、ツール出力の扱い方を制御します。
|
||||
|
||||
- `"run_llm_again"`: デフォルトです。ツールが実行され、LLM がその結果を処理して最終応答を生成します。
|
||||
- `"stop_on_first_tool"`: 最初のツール呼び出しの出力が、追加の LLM 処理なしで最終応答として使用されます。
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner, function_tool, ModelSettings
|
||||
|
||||
@function_tool
|
||||
def get_weather(city: str) -> str:
|
||||
"""Returns weather info for the specified city."""
|
||||
return f"The weather in {city} is sunny"
|
||||
|
||||
agent = Agent(
|
||||
name="Weather Agent",
|
||||
instructions="Retrieve weather details.",
|
||||
tools=[get_weather],
|
||||
tool_use_behavior="stop_on_first_tool"
|
||||
)
|
||||
```
|
||||
|
||||
- `StopAtTools(stop_at_tool_names=[...])`: 指定されたツールのいずれかが呼び出された場合に停止し、その出力を最終応答として使用します。
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner, function_tool
|
||||
from agents.agent import StopAtTools
|
||||
|
||||
@function_tool
|
||||
def get_weather(city: str) -> str:
|
||||
"""Returns weather info for the specified city."""
|
||||
return f"The weather in {city} is sunny"
|
||||
|
||||
@function_tool
|
||||
def sum_numbers(a: int, b: int) -> int:
|
||||
"""Adds two numbers."""
|
||||
return a + b
|
||||
|
||||
agent = Agent(
|
||||
name="Stop At Stock Agent",
|
||||
instructions="Get weather or sum numbers.",
|
||||
tools=[get_weather, sum_numbers],
|
||||
tool_use_behavior=StopAtTools(stop_at_tool_names=["get_weather"])
|
||||
)
|
||||
```
|
||||
|
||||
- `ToolsToFinalOutputFunction`: ツールの結果を処理し、停止するか、LLM による処理を続行するかを決定するカスタム関数です。
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner, function_tool, FunctionToolResult, RunContextWrapper
|
||||
from agents.agent import ToolsToFinalOutputResult
|
||||
from typing import List, Any
|
||||
|
||||
@function_tool
|
||||
def get_weather(city: str) -> str:
|
||||
"""Returns weather info for the specified city."""
|
||||
return f"The weather in {city} is sunny"
|
||||
|
||||
def custom_tool_handler(
|
||||
context: RunContextWrapper[Any],
|
||||
tool_results: List[FunctionToolResult]
|
||||
) -> ToolsToFinalOutputResult:
|
||||
"""Processes tool results to decide final output."""
|
||||
for result in tool_results:
|
||||
if result.output and "sunny" in result.output:
|
||||
return ToolsToFinalOutputResult(
|
||||
is_final_output=True,
|
||||
final_output=f"Final weather: {result.output}"
|
||||
)
|
||||
return ToolsToFinalOutputResult(
|
||||
is_final_output=False,
|
||||
final_output=None
|
||||
)
|
||||
|
||||
agent = Agent(
|
||||
name="Weather Agent",
|
||||
instructions="Retrieve weather details.",
|
||||
tools=[get_weather],
|
||||
tool_use_behavior=custom_tool_handler
|
||||
)
|
||||
```
|
||||
|
||||
!!! note
|
||||
|
||||
無限ループを防ぐため、フレームワークはツール呼び出し後に自動的に `tool_choice` を "auto" にリセットします。この挙動は [`agent.reset_tool_choice`][agents.agent.Agent.reset_tool_choice] で設定可能です。無限ループは、ツールの execution results が LLM に送信され、`tool_choice` のために再度ツール呼び出しが発生し、これが繰り返されることで発生します。
|
||||
|
||||
ツール呼び出し後にエージェントを完全に停止させたい場合(auto モードで継続させたくない場合)は、[`Agent.tool_use_behavior="stop_on_first_tool"`] を設定できます。これにより、ツールの出力がそのまま最終応答として使用され、以降の LLM 処理は行われません。
|
||||
無限ループを防ぐため、フレームワークはツール呼び出し後に `tool_choice` を自動的に "auto" にリセットします。この動作は [`agent.reset_tool_choice`][agents.agent.Agent.reset_tool_choice] で設定できます。無限ループが発生するのは、ツールの結果が LLM に送信され、その後 `tool_choice` のために LLM が別のツール呼び出しを生成し、これが際限なく繰り返されるためです。
|
||||
+94
-15
@@ -1,8 +1,24 @@
|
||||
# SDK の設定
|
||||
---
|
||||
search:
|
||||
exclude: true
|
||||
---
|
||||
# 設定
|
||||
|
||||
このページでは、デフォルトの OpenAI キーまたはクライアント、デフォルトの OpenAI API の形式、トレーシングエクスポートのデフォルト、ログ記録の動作など、アプリケーション起動時に通常一度だけ設定する SDK 全体のデフォルトについて説明します。
|
||||
|
||||
これらのデフォルトはサンドボックスベースのワークフローにも適用されますが、サンドボックスワークスペース、サンドボックスクライアント、セッション再利用は別途設定します。
|
||||
|
||||
代わりに特定のエージェントまたは実行を設定する必要がある場合は、まず次を参照してください:
|
||||
|
||||
- [エージェント](agents.md): 通常の `Agent` における instructions、tools、出力型、ハンドオフ、ガードレールについて。
|
||||
- [エージェントの実行](running_agents.md): `RunConfig`、セッション、会話状態のオプションについて。
|
||||
- [サンドボックスエージェント](sandbox/guide.md): `SandboxRunConfig`、マニフェスト、機能、サンドボックスクライアント固有のワークスペース設定について。
|
||||
- [モデル](models/index.md): モデル選択とプロバイダー設定について。
|
||||
- [トレーシング](tracing.md): 実行ごとのトレーシングメタデータとカスタムトレースプロセッサーについて。
|
||||
|
||||
## API キーとクライアント
|
||||
|
||||
デフォルトでは、SDK はインポート時に LLM リクエストやトレーシングのために `OPENAI_API_KEY` 環境変数を探します。アプリの起動前にこの環境変数を設定できない場合は、[set_default_openai_key()][agents.set_default_openai_key] 関数を使ってキーを設定できます。
|
||||
デフォルトでは、SDK は LLM リクエストとトレーシングに `OPENAI_API_KEY` 環境変数を使用します。このキーは、SDK が最初に OpenAI クライアントを作成するときに解決されます(遅延初期化)。そのため、最初のモデル呼び出しの前に環境変数を設定してください。アプリの起動前にその環境変数を設定できない場合は、[set_default_openai_key()][agents.set_default_openai_key] 関数を使用してキーを設定できます。
|
||||
|
||||
```python
|
||||
from agents import set_default_openai_key
|
||||
@@ -10,7 +26,7 @@ from agents import set_default_openai_key
|
||||
set_default_openai_key("sk-...")
|
||||
```
|
||||
|
||||
また、使用する OpenAI クライアントを設定することも可能です。デフォルトでは、SDK は環境変数または上記で設定したデフォルトキーを使って `AsyncOpenAI` インスタンスを作成します。これを変更したい場合は、[set_default_openai_client()][agents.set_default_openai_client] 関数を利用してください。
|
||||
また、使用する OpenAI クライアントを設定することもできます。デフォルトでは、SDK は環境変数の API キー、または上記で設定したデフォルトキーを使用して、`AsyncOpenAI` インスタンスを作成します。これは [set_default_openai_client()][agents.set_default_openai_client] 関数を使用して変更できます。
|
||||
|
||||
```python
|
||||
from openai import AsyncOpenAI
|
||||
@@ -20,7 +36,14 @@ custom_client = AsyncOpenAI(base_url="...", api_key="...")
|
||||
set_default_openai_client(custom_client)
|
||||
```
|
||||
|
||||
さらに、使用する OpenAI API をカスタマイズすることもできます。デフォルトでは OpenAI Responses API を使用していますが、[set_default_openai_api()][agents.set_default_openai_api] 関数を使って Chat Completions API を利用するように上書きできます。
|
||||
環境変数ベースのエンドポイント設定を使用したい場合、デフォルトの OpenAI プロバイダーは `OPENAI_BASE_URL` も読み取ります。Responses websocket トランスポートを有効にすると、websocket の `/responses` エンドポイント用に `OPENAI_WEBSOCKET_BASE_URL` も読み取ります。
|
||||
|
||||
```bash
|
||||
export OPENAI_BASE_URL="https://your-openai-compatible-endpoint.example/v1"
|
||||
export OPENAI_WEBSOCKET_BASE_URL="wss://your-openai-compatible-endpoint.example/v1"
|
||||
```
|
||||
|
||||
最後に、使用する OpenAI API もカスタマイズできます。デフォルトでは OpenAI Responses API を使用します。[set_default_openai_api()][agents.set_default_openai_api] 関数を使用すると、これを Chat Completions API に上書きできます。
|
||||
|
||||
```python
|
||||
from agents import set_default_openai_api
|
||||
@@ -30,7 +53,7 @@ set_default_openai_api("chat_completions")
|
||||
|
||||
## トレーシング
|
||||
|
||||
トレーシングはデフォルトで有効になっています。デフォルトでは、上記のセクションで説明した OpenAI API キー(環境変数または設定したデフォルトキー)を使用します。トレーシング専用の API キーを設定したい場合は、[`set_tracing_export_api_key`][agents.set_tracing_export_api_key] 関数を利用してください。
|
||||
トレーシングはデフォルトで有効です。デフォルトでは、上のセクションで説明したモデルリクエストと同じ OpenAI API キー(つまり、環境変数または設定したデフォルトキー)を使用します。トレーシングに使用する API キーは、[`set_tracing_export_api_key`][agents.set_tracing_export_api_key] 関数を使用して個別に設定できます。
|
||||
|
||||
```python
|
||||
from agents import set_tracing_export_api_key
|
||||
@@ -38,7 +61,41 @@ from agents import set_tracing_export_api_key
|
||||
set_tracing_export_api_key("sk-...")
|
||||
```
|
||||
|
||||
また、[`set_tracing_disabled()`][agents.set_tracing_disabled] 関数を使ってトレーシングを完全に無効化することもできます。
|
||||
モデルのトラフィックにはあるキーまたはクライアントを使用し、トレーシングには別の OpenAI キーを使用したい場合は、デフォルトキーまたはクライアントを設定するときに `use_for_tracing=False` を渡してから、トレーシングを別途設定します。カスタムクライアントを使用していない場合は、[`set_default_openai_key()`][agents.set_default_openai_key] でも同じパターンを使用できます。
|
||||
|
||||
```python
|
||||
from openai import AsyncOpenAI
|
||||
from agents import (
|
||||
set_default_openai_client,
|
||||
set_tracing_export_api_key,
|
||||
)
|
||||
|
||||
custom_client = AsyncOpenAI(base_url="https://your-openai-compatible-endpoint.example/v1", api_key="provider-key")
|
||||
set_default_openai_client(custom_client, use_for_tracing=False)
|
||||
|
||||
set_tracing_export_api_key("sk-tracing")
|
||||
```
|
||||
|
||||
デフォルトのエクスポーターを使用する際に、トレースを特定の組織またはプロジェクトに関連付ける必要がある場合は、アプリの起動前にこれらの環境変数を設定してください:
|
||||
|
||||
```bash
|
||||
export OPENAI_ORG_ID="org_..."
|
||||
export OPENAI_PROJECT_ID="proj_..."
|
||||
```
|
||||
|
||||
グローバルエクスポーターを変更せずに、実行ごとにトレーシング API キーを設定することもできます。
|
||||
|
||||
```python
|
||||
from agents import Runner, RunConfig
|
||||
|
||||
await Runner.run(
|
||||
agent,
|
||||
input="Hello",
|
||||
run_config=RunConfig(tracing={"api_key": "sk-tracing-123"}),
|
||||
)
|
||||
```
|
||||
|
||||
[`set_tracing_disabled()`][agents.set_tracing_disabled] 関数を使用して、トレーシングを完全に無効化することもできます。
|
||||
|
||||
```python
|
||||
from agents import set_tracing_disabled
|
||||
@@ -46,11 +103,31 @@ from agents import set_tracing_disabled
|
||||
set_tracing_disabled(True)
|
||||
```
|
||||
|
||||
トレーシングは有効のまま、機密性がある可能性のある入力/出力をトレースペイロードから除外したい場合は、[`RunConfig.trace_include_sensitive_data`][agents.run.RunConfig.trace_include_sensitive_data] を `False` に設定します:
|
||||
|
||||
```python
|
||||
from agents import Runner, RunConfig
|
||||
|
||||
await Runner.run(
|
||||
agent,
|
||||
input="Hello",
|
||||
run_config=RunConfig(trace_include_sensitive_data=False),
|
||||
)
|
||||
```
|
||||
|
||||
アプリの起動前にこの環境変数を設定することで、コードを変更せずにデフォルトを変更することもできます:
|
||||
|
||||
```bash
|
||||
export OPENAI_AGENTS_TRACE_INCLUDE_SENSITIVE_DATA=0
|
||||
```
|
||||
|
||||
トレーシング制御の詳細は、[トレーシングガイド](tracing.md) を参照してください。
|
||||
|
||||
## デバッグログ
|
||||
|
||||
SDK には、ハンドラーが設定されていない 2 つの Python ロガーがあります。デフォルトでは、警告やエラーは `stdout` に送信されますが、それ以外のログは抑制されます。
|
||||
SDK は 2 つの Python ロガー(`openai.agents` と `openai.agents.tracing`)を定義しますが、デフォルトではハンドラーをアタッチしません。ログはアプリケーションの Python ロギング設定に従います。
|
||||
|
||||
詳細なログ出力を有効にするには、[`enable_verbose_stdout_logging()`][agents.enable_verbose_stdout_logging] 関数を使用してください。
|
||||
詳細なログ出力を有効にするには、[`enable_verbose_stdout_logging()`][agents.enable_verbose_stdout_logging] 関数を使用します。
|
||||
|
||||
```python
|
||||
from agents import enable_verbose_stdout_logging
|
||||
@@ -58,7 +135,7 @@ from agents import enable_verbose_stdout_logging
|
||||
enable_verbose_stdout_logging()
|
||||
```
|
||||
|
||||
また、ハンドラーやフィルター、フォーマッターなどを追加してログをカスタマイズすることも可能です。詳細は [Python ロギングガイド](https://docs.python.org/3/howto/logging.html) をご覧ください。
|
||||
または、ハンドラー、フィルター、フォーマッターなどを追加してログをカスタマイズできます。詳細は [Python ロギングガイド](https://docs.python.org/3/howto/logging.html) で確認できます。
|
||||
|
||||
```python
|
||||
import logging
|
||||
@@ -77,18 +154,20 @@ logger.setLevel(logging.WARNING)
|
||||
logger.addHandler(logging.StreamHandler())
|
||||
```
|
||||
|
||||
### ログ内の機微なデータ
|
||||
### ログ内の機密データ
|
||||
|
||||
一部のログには機微なデータ(たとえば ユーザー データ)が含まれる場合があります。これらのデータのログ出力を無効にしたい場合は、以下の環境変数を設定してください。
|
||||
一部のログには機密データ(たとえば、ユーザーデータ)が含まれる場合があります。
|
||||
|
||||
LLM の入力および出力のログ出力を無効にするには:
|
||||
デフォルトでは、SDK は LLM の入力/出力やツールの入力/出力を **ログに記録しません** 。これらの保護は次で制御されます:
|
||||
|
||||
```bash
|
||||
export OPENAI_AGENTS_DONT_LOG_MODEL_DATA=1
|
||||
OPENAI_AGENTS_DONT_LOG_MODEL_DATA=1
|
||||
OPENAI_AGENTS_DONT_LOG_TOOL_DATA=1
|
||||
```
|
||||
|
||||
ツールの入力および出力のログ出力を無効にするには:
|
||||
デバッグのためにこのデータを一時的に含める必要がある場合は、アプリの起動前にいずれかの変数を `0`(または `false`)に設定します:
|
||||
|
||||
```bash
|
||||
export OPENAI_AGENTS_DONT_LOG_TOOL_DATA=1
|
||||
export OPENAI_AGENTS_DONT_LOG_MODEL_DATA=0
|
||||
export OPENAI_AGENTS_DONT_LOG_TOOL_DATA=0
|
||||
```
|
||||
+94
-23
@@ -1,29 +1,52 @@
|
||||
---
|
||||
search:
|
||||
exclude: true
|
||||
---
|
||||
# コンテキスト管理
|
||||
|
||||
コンテキストは多義的な用語です。主に関心を持つべきコンテキストには、次の 2 つの大きなクラスがあります。
|
||||
コンテキストは多義的な用語です。考慮すべきコンテキストには、主に 2 つの種類があります:
|
||||
|
||||
1. コード内でローカルに利用可能なコンテキスト:これは、ツール関数の実行時や `on_handoff` のようなコールバック、ライフサイクルフックなどで必要となるデータや依存関係です。
|
||||
2. LLM に利用可能なコンテキスト:これは、LLM がレスポンスを生成する際に参照できるデータです。
|
||||
1. コードからローカルに利用できるコンテキスト: これは、ツール関数の実行時、`on_handoff` のようなコールバック内、ライフサイクルフック内などで必要になる可能性のあるデータや依存関係です。
|
||||
2. LLM が利用できるコンテキスト: これは、LLM が応答を生成するときに参照するデータです。
|
||||
|
||||
## ローカルコンテキスト
|
||||
|
||||
これは [`RunContextWrapper`][agents.run_context.RunContextWrapper] クラスおよびその中の [`context`][agents.run_context.RunContextWrapper.context] プロパティによって表現されます。仕組みは以下の通りです。
|
||||
これは [`RunContextWrapper`][agents.run_context.RunContextWrapper] クラス、およびその中の [`context`][agents.run_context.RunContextWrapper.context] プロパティで表現されます。仕組みは次のとおりです:
|
||||
|
||||
1. 任意の Python オブジェクトを作成します。一般的なパターンとしては、dataclass や Pydantic オブジェクトを使います。
|
||||
2. そのオブジェクトを各種 run メソッド(例:`Runner.run(..., **context=whatever**))`)に渡します。
|
||||
3. すべてのツール呼び出しやライフサイクルフックなどには、ラッパーオブジェクト `RunContextWrapper[T]` が渡されます。ここで `T` はコンテキストオブジェクトの型を表し、`wrapper.context` からアクセスできます。
|
||||
1. 任意の Python オブジェクトを作成します。一般的なパターンは dataclass や Pydantic オブジェクトを使用することです。
|
||||
2. そのオブジェクトを各種実行メソッドに渡します (例: `Runner.run(..., context=whatever)`)。
|
||||
3. すべてのツール呼び出し、ライフサイクルフックなどには、ラッパーオブジェクト `RunContextWrapper[T]` が渡されます。ここで `T` はコンテキストオブジェクトの型を表し、`wrapper.context` を通じてアクセスできます。
|
||||
|
||||
**最も重要**な注意点:特定のエージェント実行において、すべてのエージェント、ツール関数、ライフサイクルなどは、同じ _型_ のコンテキストを使用する必要があります。
|
||||
一部のランタイム固有のコールバックでは、SDK はより特殊化された `RunContextWrapper[T]` のサブクラスを渡す場合があります。たとえば、関数ツールのライフサイクルフックは通常 `ToolContext` を受け取り、これは `tool_call_id`、`tool_name`、`tool_arguments` などのツール呼び出しメタデータも公開します。
|
||||
|
||||
コンテキストは以下のような用途で利用できます。
|
||||
認識すべき **最も重要な** 点は、あるエージェント実行におけるすべてのエージェント、ツール関数、ライフサイクルなどが、同じ _型_ のコンテキストを使用しなければならないということです。
|
||||
|
||||
- 実行時のコンテキストデータ(例:ユーザー名/uid やユーザーに関するその他の情報など)
|
||||
- 依存関係(例:ロガーオブジェクト、データフェッチャーなど)
|
||||
コンテキストは、たとえば次の用途に使用できます:
|
||||
|
||||
- 実行時のコンテキストデータ (例: ユーザー名 / uid や、ユーザーに関するその他の情報)
|
||||
- 依存関係 (例: ロガーオブジェクト、データ取得器など)
|
||||
- ヘルパー関数
|
||||
|
||||
!!! danger "注意"
|
||||
!!! danger "注記"
|
||||
|
||||
コンテキストオブジェクトは **LLM には送信されません**。これは純粋にローカルなオブジェクトであり、読み書きやメソッド呼び出しが可能です。
|
||||
コンテキストオブジェクトは LLM に **送信されません**。これは完全にローカルなオブジェクトであり、読み取り、書き込み、メソッドの呼び出しができます。
|
||||
|
||||
1 回の実行内では、派生したラッパーは同じ基盤となるアプリコンテキスト、承認状態、使用状況の追跡を共有します。ネストされた [`Agent.as_tool()`][agents.agent.Agent.as_tool] の実行では異なる `tool_input` を付加する場合がありますが、デフォルトではアプリ状態の分離コピーは取得しません。
|
||||
|
||||
### `RunContextWrapper` の公開内容
|
||||
|
||||
[`RunContextWrapper`][agents.run_context.RunContextWrapper] は、アプリで定義したコンテキストオブジェクトを包むラッパーです。実際には、ほとんどの場合、次のものを使用します:
|
||||
|
||||
- [`wrapper.context`][agents.run_context.RunContextWrapper.context]: 独自の可変なアプリ状態と依存関係に使用します。
|
||||
- [`wrapper.usage`][agents.run_context.RunContextWrapper.usage]: 現在の実行全体で集計されたリクエストおよびトークン使用量に使用します。
|
||||
- [`wrapper.tool_input`][agents.run_context.RunContextWrapper.tool_input]: 現在の実行が [`Agent.as_tool()`][agents.agent.Agent.as_tool] の内部で実行されている場合の構造化入力に使用します。
|
||||
- [`wrapper.approve_tool(...)`][agents.run_context.RunContextWrapper.approve_tool] / [`wrapper.reject_tool(...)`][agents.run_context.RunContextWrapper.reject_tool]: 承認状態をプログラムで更新する必要がある場合に使用します。
|
||||
|
||||
アプリで定義したオブジェクトは `wrapper.context` のみです。その他のフィールドは SDK が管理するランタイムメタデータです。
|
||||
|
||||
後で human-in-the-loop や耐久ジョブワークフローのために [`RunState`][agents.run_state.RunState] をシリアライズする場合、そのランタイムメタデータは状態とともに保存されます。シリアライズされた状態を永続化または送信する予定がある場合、[`RunContextWrapper.context`][agents.run_context.RunContextWrapper.context] にシークレットを入れないでください。
|
||||
|
||||
会話状態は別の関心事です。ターンをどのように引き継ぐかに応じて、`result.to_input_list()`、`session`、`conversation_id`、または `previous_response_id` を使用してください。その判断については、[実行結果](results.md)、[エージェントの実行](running_agents.md)、[セッション](sessions/index.md) を参照してください。
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
@@ -38,7 +61,8 @@ class UserInfo: # (1)!
|
||||
|
||||
@function_tool
|
||||
async def fetch_user_age(wrapper: RunContextWrapper[UserInfo]) -> str: # (2)!
|
||||
return f"User {wrapper.context.name} is 47 years old"
|
||||
"""Fetch the age of the user. Call this function to get user's age information."""
|
||||
return f"The user {wrapper.context.name} is 47 years old"
|
||||
|
||||
async def main():
|
||||
user_info = UserInfo(name="John", uid=123)
|
||||
@@ -61,17 +85,64 @@ if __name__ == "__main__":
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
1. これはコンテキストオブジェクトです。ここでは dataclass を使用していますが、任意の型を利用できます。
|
||||
2. これはツールです。`RunContextWrapper[UserInfo]` を受け取っていることが分かります。ツールの実装はコンテキストから値を読み取ります。
|
||||
3. エージェントにはジェネリック型 `UserInfo` を指定しています。これにより、型チェッカーがエラーを検出できます(例えば、異なるコンテキスト型を受け取るツールを渡そうとした場合など)。
|
||||
1. これがコンテキストオブジェクトです。ここでは dataclass を使用していますが、任意の型を使用できます。
|
||||
2. これはツールです。`RunContextWrapper[UserInfo]` を受け取っていることがわかります。ツール実装はコンテキストから読み取ります。
|
||||
3. 型チェッカーがエラーを検出できるように、エージェントにジェネリック `UserInfo` を指定します (たとえば、異なるコンテキスト型を受け取るツールを渡そうとした場合)。
|
||||
4. コンテキストは `run` 関数に渡されます。
|
||||
5. エージェントは正しくツールを呼び出し、年齢を取得します。
|
||||
|
||||
## エージェント/LLM コンテキスト
|
||||
---
|
||||
|
||||
LLM が呼び出される際、**唯一** 参照できるデータは会話履歴からのものです。つまり、LLM に新しいデータを利用させたい場合は、そのデータを履歴に含める必要があります。これを実現する方法はいくつかあります。
|
||||
### 高度な内容: `ToolContext`
|
||||
|
||||
1. エージェントの `instructions` に追加する。この方法は「システムプロンプト」や「開発者メッセージ」とも呼ばれます。システムプロンプトは静的な文字列でも、コンテキストを受け取って文字列を出力する動的な関数でも構いません。たとえば、ユーザー名や現在の日付など、常に有用な情報に適しています。
|
||||
2. `Runner.run` 関数を呼び出す際に `input` に追加する。この方法は `instructions` と似ていますが、[chain of command](https://cdn.openai.com/spec/model-spec-2024-05-08.html#follow-the-chain-of-command) の下位メッセージとして追加できます。
|
||||
3. 関数ツールを通じて公開する。この方法は _オンデマンド_ のコンテキストに適しています。LLM が必要なタイミングでツールを呼び出し、データを取得できます。
|
||||
4. リトリーバルや Web 検索を利用する。これらはファイルやデータベース(リトリーバル)、または Web(Web 検索)から関連データを取得できる特別なツールです。関連するコンテキストデータに基づいたレスポンスを「グラウンディング」するのに役立ちます。
|
||||
場合によっては、実行中のツールに関する追加メタデータ (名前、呼び出し ID、生の引数文字列など) にアクセスしたいことがあります。
|
||||
この場合、`RunContextWrapper` を拡張する [`ToolContext`][agents.tool_context.ToolContext] クラスを使用できます。
|
||||
|
||||
```python
|
||||
from typing import Annotated
|
||||
from pydantic import BaseModel, Field
|
||||
from agents import Agent, Runner, function_tool
|
||||
from agents.tool_context import ToolContext
|
||||
|
||||
class WeatherContext(BaseModel):
|
||||
user_id: str
|
||||
|
||||
class Weather(BaseModel):
|
||||
city: str = Field(description="The city name")
|
||||
temperature_range: str = Field(description="The temperature range in Celsius")
|
||||
conditions: str = Field(description="The weather conditions")
|
||||
|
||||
@function_tool
|
||||
def get_weather(ctx: ToolContext[WeatherContext], city: Annotated[str, "The city to get the weather for"]) -> Weather:
|
||||
print(f"[debug] Tool context: (name: {ctx.tool_name}, call_id: {ctx.tool_call_id}, args: {ctx.tool_arguments})")
|
||||
return Weather(city=city, temperature_range="14-20C", conditions="Sunny with wind.")
|
||||
|
||||
agent = Agent(
|
||||
name="Weather Agent",
|
||||
instructions="You are a helpful agent that can tell the weather of a given city.",
|
||||
tools=[get_weather],
|
||||
)
|
||||
```
|
||||
|
||||
`ToolContext` は `RunContextWrapper` と同じ `.context` プロパティを提供し、
|
||||
現在のツール呼び出しに固有の追加フィールドも提供します:
|
||||
|
||||
- `tool_name` – 呼び出されているツールの名前
|
||||
- `tool_call_id` – このツール呼び出しの一意の識別子
|
||||
- `tool_arguments` – ツールに渡された生の引数文字列
|
||||
- `tool_namespace` – ツールが `tool_namespace()` または別の名前空間付きサーフェスを通じて読み込まれた場合の、ツール呼び出しに対する Responses 名前空間
|
||||
- `qualified_tool_name` – 名前空間が利用できる場合に、その名前空間で修飾されたツール名
|
||||
|
||||
実行中にツールレベルのメタデータが必要な場合は、`ToolContext` を使用してください。
|
||||
エージェントとツール間で一般的なコンテキスト共有を行うには、`RunContextWrapper` のままで十分です。`ToolContext` は `RunContextWrapper` を拡張しているため、ネストされた `Agent.as_tool()` 実行が構造化入力を提供した場合には `.tool_input` も公開できます。
|
||||
|
||||
---
|
||||
|
||||
## エージェント / LLM コンテキスト
|
||||
|
||||
LLM が呼び出されるとき、その LLM が参照できる **唯一の** データは会話履歴に含まれるものです。つまり、新しいデータを LLM に利用可能にしたい場合は、その履歴内で利用可能になるような方法で行う必要があります。これにはいくつかの方法があります:
|
||||
|
||||
1. エージェントの `instructions` に追加できます。これは「システムプロンプト」または「開発者メッセージ」とも呼ばれます。システムプロンプトは静的文字列にも、コンテキストを受け取って文字列を出力する動的関数にもできます。これは、常に役立つ情報 (たとえば、ユーザーの名前や現在の日付) に対する一般的な手法です。
|
||||
2. `Runner.run` 関数を呼び出すときに `input` に追加します。これは `instructions` の手法に似ていますが、[指揮系統](https://cdn.openai.com/spec/model-spec-2024-05-08.html#follow-the-chain-of-command) においてより下位のメッセージにできます。
|
||||
3. 関数ツールを介して公開します。これは _オンデマンド_ のコンテキストに便利です。LLM がデータを必要とするタイミングを判断し、そのデータを取得するためにツールを呼び出せます。
|
||||
4. リトリーバルまたは Web 検索を使用します。これらは、ファイルやデータベースから関連データを取得する (リトリーバル)、または Web から取得する (Web 検索) ことができる特殊なツールです。これは、関連するコンテキストデータに基づいて応答を「グラウンディング」するのに便利です。
|
||||
+127
-25
@@ -1,40 +1,142 @@
|
||||
---
|
||||
search:
|
||||
exclude: true
|
||||
---
|
||||
# コード例
|
||||
|
||||
SDK のさまざまなサンプル実装については、[リポジトリ](https://github.com/openai/openai-agents-python/tree/main/examples) のコード例セクションをご覧ください。これらのコード例は、異なるパターンや機能を示すいくつかのカテゴリーに整理されています。
|
||||
SDK のさまざまなサンプル実装は、[リポジトリ](https://github.com/openai/openai-agents-python/tree/main/examples)の examples セクションで確認できます。これらのコード例は、さまざまなパターンや機能を示す複数のカテゴリーに整理されています。
|
||||
|
||||
## カテゴリー
|
||||
|
||||
- **[agent_patterns](https://github.com/openai/openai-agents-python/tree/main/examples/agent_patterns):**
|
||||
このカテゴリーのコード例では、よく使われるエージェント設計パターンを紹介しています。
|
||||
- **[agent_patterns](https://github.com/openai/openai-agents-python/tree/main/examples/agent_patterns):**
|
||||
このカテゴリーのコード例は、次のような一般的なエージェント設計パターンを示します。
|
||||
|
||||
- 決定論的なワークフロー
|
||||
- ツールとしてのエージェント
|
||||
- エージェントの並列実行
|
||||
- 決定論的ワークフロー
|
||||
- Agents as tools
|
||||
- ストリーミングイベントを伴う Agents as tools (`examples/agent_patterns/agents_as_tools_streaming.py`)
|
||||
- 構造化入力パラメーターを伴う Agents as tools (`examples/agent_patterns/agents_as_tools_structured.py`)
|
||||
- エージェントの並列実行
|
||||
- 条件付きツール使用
|
||||
- 異なる動作でツール使用を強制 (`examples/agent_patterns/forcing_tool_use.py`)
|
||||
- 入出力ガードレール
|
||||
- ジャッジとしての LLM
|
||||
- ルーティング
|
||||
- ストリーミングガードレール
|
||||
- ツール承認と状態シリアライズを伴う人間参加型 (`examples/agent_patterns/human_in_the_loop.py`)
|
||||
- ストリーミングを伴う人間参加型 (`examples/agent_patterns/human_in_the_loop_stream.py`)
|
||||
- 承認フロー用のカスタム拒否メッセージ (`examples/agent_patterns/human_in_the_loop_custom_rejection.py`)
|
||||
|
||||
- **[basic](https://github.com/openai/openai-agents-python/tree/main/examples/basic):**
|
||||
これらのコード例では、SDK の基本的な機能を紹介しています。
|
||||
- **[basic](https://github.com/openai/openai-agents-python/tree/main/examples/basic):**
|
||||
これらのコード例では、SDK の基本的な機能を紹介します。たとえば、次のようなものです。
|
||||
|
||||
- 動的なシステムプロンプト
|
||||
- ストリーミング出力
|
||||
- ライフサイクルイベント
|
||||
- Hello world の例 (デフォルトモデル、GPT-5、オープンウェイトモデル)
|
||||
- エージェントライフサイクル管理
|
||||
- 実行フックとエージェントフックのライフサイクル例 (`examples/basic/lifecycle_example.py`)
|
||||
- 動的システムプロンプト
|
||||
- 基本的なツール使用 (`examples/basic/tools.py`)
|
||||
- ツール入出力ガードレール (`examples/basic/tool_guardrails.py`)
|
||||
- 画像ツール出力 (`examples/basic/image_tool_output.py`)
|
||||
- ストリーミング出力 (テキスト、アイテム、関数呼び出し引数)
|
||||
- ターン間で共有セッションヘルパーを使用する Responses WebSocket トランスポート (`examples/basic/stream_ws.py`)
|
||||
- プロンプトテンプレート
|
||||
- ファイル処理 (ローカルとリモート、画像と PDF)
|
||||
- 使用状況の追跡
|
||||
- Runner 管理のリトライ設定 (`examples/basic/retry.py`)
|
||||
- サードパーティアダプターを介した Runner 管理のリトライ (`examples/basic/retry_litellm.py`)
|
||||
- 非厳密な出力型
|
||||
- 以前のレスポンス ID の使用
|
||||
|
||||
- **[tool examples](https://github.com/openai/openai-agents-python/tree/main/examples/tools):**
|
||||
OpenAI がホストするツール(Web 検索やファイル検索など)の実装方法や、それらをエージェントに統合する方法を学べます。
|
||||
- **[customer_service](https://github.com/openai/openai-agents-python/tree/main/examples/customer_service):**
|
||||
航空会社向けのカスタマーサービスシステムの例です。
|
||||
|
||||
- **[model providers](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers):**
|
||||
OpenAI 以外のモデルを SDK で利用する方法を紹介しています。
|
||||
- **[financial_research_agent](https://github.com/openai/openai-agents-python/tree/main/examples/financial_research_agent):**
|
||||
金融データ分析向けのエージェントとツールを使った、構造化されたリサーチワークフローを示す金融リサーチエージェントです。
|
||||
|
||||
- **[handoffs](https://github.com/openai/openai-agents-python/tree/main/examples/handoffs):**
|
||||
エージェントのハンドオフの実践的なコード例をご覧いただけます。
|
||||
- **[handoffs](https://github.com/openai/openai-agents-python/tree/main/examples/handoffs):**
|
||||
メッセージフィルタリングを伴うエージェントのハンドオフの実用的なコード例です。以下を含みます:
|
||||
|
||||
- **[mcp](https://github.com/openai/openai-agents-python/tree/main/examples/mcp):**
|
||||
Model context protocol (MCP) を使ったエージェントの構築方法を学べます。
|
||||
- メッセージフィルターの例 (`examples/handoffs/message_filter.py`)
|
||||
- ストリーミングを伴うメッセージフィルター (`examples/handoffs/message_filter_streaming.py`)
|
||||
|
||||
- **[customer_service](https://github.com/openai/openai-agents-python/tree/main/examples/customer_service)** および **[research_bot](https://github.com/openai/openai-agents-python/tree/main/examples/research_bot):**
|
||||
実際のユースケースを示す、より発展的な 2 つのコード例です。
|
||||
- **[hosted_mcp](https://github.com/openai/openai-agents-python/tree/main/examples/hosted_mcp):**
|
||||
OpenAI Responses API でホスト型 MCP (Model Context Protocol) を使用する方法を示すコード例です。以下を含みます:
|
||||
|
||||
- **customer_service**: 航空会社向けカスタマーサービスシステムの例。
|
||||
- **research_bot**: シンプルなディープリサーチクローン。
|
||||
- 承認なしのシンプルなホスト型 MCP (`examples/hosted_mcp/simple.py`)
|
||||
- Google Calendar などの MCP コネクター (`examples/hosted_mcp/connectors.py`)
|
||||
- 中断ベースの承認を伴う人間参加型 (`examples/hosted_mcp/human_in_the_loop.py`)
|
||||
- MCP ツール呼び出しの承認時コールバック (`examples/hosted_mcp/on_approval.py`)
|
||||
|
||||
- **[voice](https://github.com/openai/openai-agents-python/tree/main/examples/voice):**
|
||||
TTS および STT モデルを利用した音声エージェントのコード例をご覧いただけます。
|
||||
- **[mcp](https://github.com/openai/openai-agents-python/tree/main/examples/mcp):**
|
||||
MCP (Model Context Protocol) でエージェントを構築する方法を学びます。以下を含みます:
|
||||
|
||||
- ファイルシステムの例
|
||||
- Git の例
|
||||
- MCP プロンプトサーバーの例
|
||||
- SSE (Server-Sent Events) の例
|
||||
- SSE リモートサーバー接続 (`examples/mcp/sse_remote_example`)
|
||||
- Streamable HTTP の例
|
||||
- Streamable HTTP リモート接続 (`examples/mcp/streamable_http_remote_example`)
|
||||
- Streamable HTTP 用のカスタム HTTP クライアントファクトリー (`examples/mcp/streamablehttp_custom_client_example`)
|
||||
- `MCPUtil.get_all_function_tools` によるすべての MCP ツールの事前取得 (`examples/mcp/get_all_mcp_tools_example`)
|
||||
- FastAPI を使用した MCPServerManager (`examples/mcp/manager_example`)
|
||||
- MCP ツールフィルタリング (`examples/mcp/tool_filter_example`)
|
||||
|
||||
- **[memory](https://github.com/openai/openai-agents-python/tree/main/examples/memory):**
|
||||
エージェント向けのさまざまなメモリ実装のコード例です。以下を含みます:
|
||||
|
||||
- SQLite セッションストレージ
|
||||
- 高度な SQLite セッションストレージ
|
||||
- Redis セッションストレージ
|
||||
- SQLAlchemy セッションストレージ
|
||||
- Dapr 状態ストアセッションストレージ
|
||||
- 暗号化セッションストレージ
|
||||
- OpenAI Conversations セッションストレージ
|
||||
- Responses 圧縮セッションストレージ
|
||||
- `ModelSettings(store=False)` を使用したステートレスな Responses 圧縮 (`examples/memory/compaction_session_stateless_example.py`)
|
||||
- ファイルバック型セッションストレージ (`examples/memory/file_session.py`)
|
||||
- 人間参加型を伴うファイルバック型セッション (`examples/memory/file_hitl_example.py`)
|
||||
- 人間参加型を伴う SQLite インメモリセッション (`examples/memory/memory_session_hitl_example.py`)
|
||||
- 人間参加型を伴う OpenAI Conversations セッション (`examples/memory/openai_session_hitl_example.py`)
|
||||
- セッションをまたぐ HITL 承認 / 拒否シナリオ (`examples/memory/hitl_session_scenario.py`)
|
||||
|
||||
- **[model_providers](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers):**
|
||||
カスタムプロバイダーやサードパーティアダプターを含め、SDK で OpenAI 以外のモデルを使用する方法を確認できます。
|
||||
|
||||
- **[realtime](https://github.com/openai/openai-agents-python/tree/main/examples/realtime):**
|
||||
SDK を使用してリアルタイム体験を構築する方法を示すコード例です。以下を含みます:
|
||||
|
||||
- 構造化テキストと画像メッセージを扱う Web アプリケーションパターン
|
||||
- コマンドライン音声ループと再生処理
|
||||
- WebSocket 経由の Twilio Media Streams 統合
|
||||
- Realtime Calls API の attach フローを使用した Twilio SIP 統合
|
||||
|
||||
- **[reasoning_content](https://github.com/openai/openai-agents-python/tree/main/examples/reasoning_content):**
|
||||
推論コンテンツの扱い方を示すコード例です。以下を含みます:
|
||||
|
||||
- Runner API での推論コンテンツ、ストリーミングと非ストリーミング (`examples/reasoning_content/runner_example.py`)
|
||||
- OpenRouter 経由の OSS モデルでの推論コンテンツ (`examples/reasoning_content/gpt_oss_stream.py`)
|
||||
- 基本的な推論コンテンツの例 (`examples/reasoning_content/main.py`)
|
||||
|
||||
- **[research_bot](https://github.com/openai/openai-agents-python/tree/main/examples/research_bot):**
|
||||
複雑なマルチエージェントリサーチワークフローを示す、シンプルなディープリサーチのクローンです。
|
||||
|
||||
- **[tools](https://github.com/openai/openai-agents-python/tree/main/examples/tools):**
|
||||
OpenAI がホストするツールや、次のような実験的な Codex ツール群の実装方法を学びます:
|
||||
|
||||
- Web 検索およびフィルター付き Web 検索
|
||||
- ファイル検索
|
||||
- Code interpreter
|
||||
- ファイル編集と承認を伴う apply patch ツール (`examples/tools/apply_patch.py`)
|
||||
- 承認コールバックを伴うシェルツール実行 (`examples/tools/shell.py`)
|
||||
- 人間参加型の中断ベース承認を伴うシェルツール (`examples/tools/shell_human_in_the_loop.py`)
|
||||
- インラインスキルを使用するホスト型コンテナーシェル (`examples/tools/container_shell_inline_skill.py`)
|
||||
- スキル参照を使用するホスト型コンテナーシェル (`examples/tools/container_shell_skill_reference.py`)
|
||||
- ローカルスキルを使用するローカルシェル (`examples/tools/local_shell_skill.py`)
|
||||
- 名前空間と遅延ツールを使用するツール検索 (`examples/tools/tool_search.py`)
|
||||
- コンピュータ操作
|
||||
- 画像生成
|
||||
- 実験的な Codex ツールワークフロー (`examples/tools/codex.py`)
|
||||
- 実験的な Codex 同一スレッドワークフロー (`examples/tools/codex_same_thread.py`)
|
||||
|
||||
- **[voice](https://github.com/openai/openai-agents-python/tree/main/examples/voice):**
|
||||
OpenAI の TTS および STT モデルを使用する音声エージェントの例をご覧ください。ストリーミング音声の例も含まれます。
|
||||
+99
-20
@@ -1,43 +1,77 @@
|
||||
---
|
||||
search:
|
||||
exclude: true
|
||||
---
|
||||
# ガードレール
|
||||
|
||||
ガードレールは、エージェントと _並行して_ 実行され、ユーザー入力のチェックやバリデーションを行うことができます。例えば、非常に賢い(そのため遅くて高価な)モデルを使ってカスタマーリクエストに対応するエージェントがあるとします。悪意のあるユーザーがモデルに数学の宿題を手伝わせるようなリクエストを送ることは避けたいでしょう。そこで、ガードレールを高速かつ安価なモデルで実行できます。ガードレールが悪意のある利用を検知した場合、即座にエラーを発生させ、高価なモデルの実行を止めて時間やコストを節約できます。
|
||||
ガードレールを使用すると、ユーザー入力とエージェント出力のチェックと検証を行えます。例えば、顧客リクエストの支援に非常に賢い(そのため遅く高コストな)モデルを使用するエージェントがあるとします。悪意あるユーザーが、そのモデルに数学の宿題を手伝わせることは望ましくありません。そのため、高速/低コストなモデルでガードレールを実行できます。ガードレールが悪用を検出した場合、ただちにエラーを発生させ、高コストなモデルの実行を防ぐことができ、時間とコストを節約できます( **ブロッキングガードレールを使用している場合です。並列ガードレールでは、ガードレールが完了する前に高コストなモデルの実行がすでに開始している可能性があります。詳細は以下の「実行モード」を参照してください** )。
|
||||
|
||||
ガードレールには 2 種類あります:
|
||||
ガードレールには 2 種類あります。
|
||||
|
||||
1. 入力ガードレール:最初のユーザー入力に対して実行されます
|
||||
2. 出力ガードレール:最終的なエージェント出力に対して実行されます
|
||||
1. 入力ガードレールは最初のユーザー入力に対して実行されます
|
||||
2. 出力ガードレールは最終的なエージェント出力に対して実行されます
|
||||
|
||||
## ワークフローの境界
|
||||
|
||||
ガードレールはエージェントやツールに付与されますが、ワークフロー内の同じ時点ですべてが実行されるわけではありません。
|
||||
|
||||
- **入力ガードレール** は、チェーン内の最初のエージェントに対してのみ実行されます。
|
||||
- **出力ガードレール** は、最終出力を生成するエージェントに対してのみ実行されます。
|
||||
- **ツールガードレール** は、すべてのカスタム関数ツール呼び出しで実行されます。入力ガードレールは実行前に、出力ガードレールは実行後に実行されます。
|
||||
|
||||
マネージャー、ハンドオフ、または委任先のスペシャリストを含むワークフローで、各カスタム関数ツール呼び出しの周囲にチェックが必要な場合は、エージェントレベルの入力/出力ガードレールだけに頼るのではなく、ツールガードレールを使用してください。
|
||||
|
||||
## 入力ガードレール
|
||||
|
||||
入力ガードレールは 3 ステップで実行されます:
|
||||
入力ガードレールは 3 ステップで実行されます。
|
||||
|
||||
1. まず、ガードレールはエージェントに渡されたものと同じ入力を受け取ります。
|
||||
2. 次に、ガードレール関数が実行され、[`GuardrailFunctionOutput`][agents.guardrail.GuardrailFunctionOutput] を生成し、それが [`InputGuardrailResult`][agents.guardrail.InputGuardrailResult] でラップされます。
|
||||
3. 最後に、[`.tripwire_triggered`][agents.guardrail.GuardrailFunctionOutput.tripwire_triggered] が true かどうかを確認します。true の場合、[`InputGuardrailTripwireTriggered`][agents.exceptions.InputGuardrailTripwireTriggered] 例外が発生し、ユーザーへの適切な対応や例外処理が可能です。
|
||||
2. 次に、ガードレール関数が実行されて [`GuardrailFunctionOutput`][agents.guardrail.GuardrailFunctionOutput] を生成し、それが [`InputGuardrailResult`][agents.guardrail.InputGuardrailResult] にラップされます
|
||||
3. 最後に、[`.tripwire_triggered`][agents.guardrail.GuardrailFunctionOutput.tripwire_triggered] が true であるかどうかを確認します。true の場合、[`InputGuardrailTripwireTriggered`][agents.exceptions.InputGuardrailTripwireTriggered] 例外が発生するため、ユーザーへ適切に応答したり、例外を処理したりできます。
|
||||
|
||||
!!! Note
|
||||
|
||||
入力ガードレールはユーザー入力に対して実行されることを想定しているため、エージェントのガードレールは *最初* のエージェントでのみ実行されます。「なぜ `guardrails` プロパティがエージェントにあり、`Runner.run` に渡さないのか?」と疑問に思うかもしれません。これは、ガードレールが実際のエージェントに関連することが多いためです。異なるエージェントごとに異なるガードレールを実行するため、コードを同じ場所にまとめておくと可読性が向上します。
|
||||
入力ガードレールはユーザー入力に対して実行されることを意図しているため、エージェントのガードレールは、そのエージェントが *最初の* エージェントである場合にのみ実行されます。なぜ `guardrails` プロパティが `Runner.run` に渡されるのではなく、エージェント上にあるのか疑問に思うかもしれません。これは、ガードレールが実際のエージェントに関連することが多いためです。エージェントごとに異なるガードレールを実行することになるため、コードを同じ場所に配置すると読みやすさの面で有用です。
|
||||
|
||||
### 実行モード
|
||||
|
||||
入力ガードレールは 2 つの実行モードをサポートします。
|
||||
|
||||
- **並列実行** (デフォルト、 `run_in_parallel=True` ): ガードレールはエージェントの実行と同時に実行されます。両方が同時に開始するため、レイテンシーの面で最良です。ただし、ガードレールが失敗した場合、キャンセルされる前にエージェントがすでにトークンを消費し、ツールを実行している可能性があります。
|
||||
|
||||
- **ブロッキング実行** ( `run_in_parallel=False` ): ガードレールはエージェントの開始 *前に* 実行され、完了します。ガードレールのトリップワイヤーがトリガーされた場合、エージェントは一切実行されないため、トークン消費とツール実行を防げます。これは、コスト最適化や、ツール呼び出しによる潜在的な副作用を避けたい場合に最適です。
|
||||
|
||||
## 出力ガードレール
|
||||
|
||||
出力ガードレールも 3 ステップで実行されます:
|
||||
出力ガードレールは 3 ステップで実行されます。
|
||||
|
||||
1. まず、ガードレールはエージェントに渡されたものと同じ入力を受け取ります。
|
||||
2. 次に、ガードレール関数が実行され、[`GuardrailFunctionOutput`][agents.guardrail.GuardrailFunctionOutput] を生成し、それが [`OutputGuardrailResult`][agents.guardrail.OutputGuardrailResult] でラップされます。
|
||||
3. 最後に、[`.tripwire_triggered`][agents.guardrail.GuardrailFunctionOutput.tripwire_triggered] が true かどうかを確認します。true の場合、[`OutputGuardrailTripwireTriggered`][agents.exceptions.OutputGuardrailTripwireTriggered] 例外が発生し、ユーザーへの適切な対応や例外処理が可能です。
|
||||
1. まず、ガードレールはエージェントによって生成された出力を受け取ります。
|
||||
2. 次に、ガードレール関数が実行されて [`GuardrailFunctionOutput`][agents.guardrail.GuardrailFunctionOutput] を生成し、それが [`OutputGuardrailResult`][agents.guardrail.OutputGuardrailResult] にラップされます
|
||||
3. 最後に、[`.tripwire_triggered`][agents.guardrail.GuardrailFunctionOutput.tripwire_triggered] が true であるかどうかを確認します。true の場合、[`OutputGuardrailTripwireTriggered`][agents.exceptions.OutputGuardrailTripwireTriggered] 例外が発生するため、ユーザーへ適切に応答したり、例外を処理したりできます。
|
||||
|
||||
!!! Note
|
||||
|
||||
出力ガードレールは最終的なエージェント出力に対して実行されることを想定しているため、エージェントのガードレールは *最後* のエージェントでのみ実行されます。入力ガードレールと同様に、ガードレールが実際のエージェントに関連することが多いため、コードを同じ場所にまとめておくと可読性が向上します。
|
||||
出力ガードレールは最終的なエージェント出力に対して実行されることを意図しているため、エージェントのガードレールは、そのエージェントが *最後の* エージェントである場合にのみ実行されます。入力ガードレールと同様に、これはガードレールが実際のエージェントに関連することが多いためです。エージェントごとに異なるガードレールを実行することになるため、コードを同じ場所に配置すると読みやすさの面で有用です。
|
||||
|
||||
出力ガードレールは常にエージェントの完了後に実行されるため、 `run_in_parallel` パラメーターはサポートしていません。
|
||||
|
||||
## ツールガードレール
|
||||
|
||||
ツールガードレールは **関数ツール** をラップし、実行前後にツール呼び出しを検証またはブロックできるようにします。ツール自体に設定され、そのツールが呼び出されるたびに実行されます。
|
||||
|
||||
- 入力ツールガードレールはツールの実行前に実行され、呼び出しのスキップ、出力のメッセージへの置き換え、またはトリップワイヤーの発生を行えます。
|
||||
- 出力ツールガードレールはツールの実行後に実行され、出力の置き換え、またはトリップワイヤーの発生を行えます。
|
||||
- ツールガードレールは、[`function_tool`][agents.tool.function_tool] で作成された関数ツールにのみ適用されます。ハンドオフは通常の関数ツールパイプラインではなく SDK のハンドオフパイプラインを通るため、ツールガードレールはハンドオフ呼び出し自体には適用されません。ホスト型ツール( `WebSearchTool` 、 `FileSearchTool` 、 `HostedMCPTool` 、 `CodeInterpreterTool` 、 `ImageGenerationTool` )や組み込み実行ツール( `ComputerTool` 、 `ShellTool` 、 `ApplyPatchTool` 、 `LocalShellTool` )もこのガードレールパイプラインを使用しません。また、[`Agent.as_tool()`][agents.agent.Agent.as_tool] は現在、ツールガードレールのオプションを直接公開していません。
|
||||
|
||||
詳細は以下のコードスニペットを参照してください。
|
||||
|
||||
## トリップワイヤー
|
||||
|
||||
入力または出力がガードレールに失敗した場合、ガードレールはトリップワイヤーでこれを通知できます。トリップワイヤーが発動したガードレールを検知した時点で、即座に `{Input,Output}GuardrailTripwireTriggered` 例外を発生させ、エージェントの実行を停止します。
|
||||
入力または出力がガードレールの検査に合格しない場合、ガードレールはトリップワイヤーでこれを通知できます。トリップワイヤーがトリガーされたガードレールを検出すると、ただちに `{Input,Output}GuardrailTripwireTriggered` 例外を発生させ、エージェント実行を停止します。
|
||||
|
||||
## ガードレールの実装
|
||||
|
||||
入力を受け取り、[`GuardrailFunctionOutput`][agents.guardrail.GuardrailFunctionOutput] を返す関数を用意する必要があります。この例では、内部でエージェントを実行することでこれを実現します。
|
||||
入力を受け取り、[`GuardrailFunctionOutput`][agents.guardrail.GuardrailFunctionOutput] を返す関数を用意する必要があります。この例では、内部でエージェントを実行することでこれを行います。
|
||||
|
||||
```python
|
||||
from pydantic import BaseModel
|
||||
@@ -90,9 +124,9 @@ async def main():
|
||||
print("Math homework guardrail tripped")
|
||||
```
|
||||
|
||||
1. このエージェントをガードレール関数内で使用します。
|
||||
2. これはエージェントの入力やコンテキストを受け取り、結果を返すガードレール関数です。
|
||||
3. ガードレールの結果に追加情報を含めることができます。
|
||||
1. このエージェントをガードレール関数で使用します。
|
||||
2. これは、エージェントの入力/コンテキストを受け取り、実行結果を返すガードレール関数です。
|
||||
3. ガードレールの実行結果に追加情報を含めることができます。
|
||||
4. これはワークフローを定義する実際のエージェントです。
|
||||
|
||||
出力ガードレールも同様です。
|
||||
@@ -150,5 +184,50 @@ async def main():
|
||||
|
||||
1. これは実際のエージェントの出力型です。
|
||||
2. これはガードレールの出力型です。
|
||||
3. これはエージェントの出力を受け取り、結果を返すガードレール関数です。
|
||||
4. これはワークフローを定義する実際のエージェントです。
|
||||
3. これは、エージェントの出力を受け取り、実行結果を返すガードレール関数です。
|
||||
4. これはワークフローを定義する実際のエージェントです。
|
||||
|
||||
最後に、ツールガードレールのコード例を示します。
|
||||
|
||||
```python
|
||||
import json
|
||||
from agents import (
|
||||
Agent,
|
||||
Runner,
|
||||
ToolGuardrailFunctionOutput,
|
||||
function_tool,
|
||||
tool_input_guardrail,
|
||||
tool_output_guardrail,
|
||||
)
|
||||
|
||||
@tool_input_guardrail
|
||||
def block_secrets(data):
|
||||
args = json.loads(data.context.tool_arguments or "{}")
|
||||
if "sk-" in json.dumps(args):
|
||||
return ToolGuardrailFunctionOutput.reject_content(
|
||||
"Remove secrets before calling this tool."
|
||||
)
|
||||
return ToolGuardrailFunctionOutput.allow()
|
||||
|
||||
|
||||
@tool_output_guardrail
|
||||
def redact_output(data):
|
||||
text = str(data.output or "")
|
||||
if "sk-" in text:
|
||||
return ToolGuardrailFunctionOutput.reject_content("Output contained sensitive data.")
|
||||
return ToolGuardrailFunctionOutput.allow()
|
||||
|
||||
|
||||
@function_tool(
|
||||
tool_input_guardrails=[block_secrets],
|
||||
tool_output_guardrails=[redact_output],
|
||||
)
|
||||
def classify_text(text: str) -> str:
|
||||
"""Classify text for internal routing."""
|
||||
return f"length:{len(text)}"
|
||||
|
||||
|
||||
agent = Agent(name="Classifier", tools=[classify_text])
|
||||
result = Runner.run_sync(agent, "hello world")
|
||||
print(result.final_output)
|
||||
```
|
||||
+61
-18
@@ -1,18 +1,24 @@
|
||||
---
|
||||
search:
|
||||
exclude: true
|
||||
---
|
||||
# ハンドオフ
|
||||
|
||||
ハンドオフは、エージェントが他のエージェントにタスクを委任できる仕組みです。これは、異なるエージェントがそれぞれ異なる分野に特化しているシナリオで特に有用です。たとえば、カスタマーサポートアプリでは、注文状況、返金、FAQ などのタスクをそれぞれ専門に扱うエージェントが存在する場合があります。
|
||||
ハンドオフにより、エージェントはタスクを別のエージェントに委任できます。これは、異なるエージェントが別々の領域に特化しているシナリオで特に有用です。たとえば、カスタマーサポートアプリには、注文状況、返金、FAQ などのタスクをそれぞれ専門的に扱うエージェントがあるかもしれません。
|
||||
|
||||
ハンドオフは LLM からはツールとして認識されます。たとえば、`Refund Agent` という名前のエージェントへのハンドオフがある場合、そのツール名は `transfer_to_refund_agent` となります。
|
||||
ハンドオフは LLM に対してツールとして表現されます。そのため、`Refund Agent` という名前のエージェントへのハンドオフがある場合、そのツールは `transfer_to_refund_agent` と呼ばれます。
|
||||
|
||||
## ハンドオフの作成
|
||||
|
||||
すべてのエージェントは [`handoffs`][agents.agent.Agent.handoffs] パラメーターを持っており、これは直接 `Agent` を指定することも、ハンドオフをカスタマイズする `Handoff` オブジェクトを指定することもできます。
|
||||
すべてのエージェントには [`handoffs`][agents.agent.Agent.handoffs] パラメーターがあり、これは `Agent` を直接受け取ることも、ハンドオフをカスタマイズする `Handoff` オブジェクトを受け取ることもできます。
|
||||
|
||||
Agents SDK で提供されている [`handoff()`][agents.handoffs.handoff] 関数を使ってハンドオフを作成できます。この関数では、ハンドオフ先のエージェントや、オプションのオーバーライドや入力フィルターを指定できます。
|
||||
単純な `Agent` インスタンスを渡した場合、それらの [`handoff_description`][agents.agent.Agent.handoff_description](設定されている場合)が既定のツール説明に追加されます。完全な `handoff()` オブジェクトを書かずに、モデルがそのハンドオフを選ぶべきタイミングを示すために使用してください。
|
||||
|
||||
Agents SDK が提供する [`handoff()`][agents.handoffs.handoff] 関数を使用して、ハンドオフを作成できます。この関数では、ハンドオフ先のエージェントに加え、省略可能なオーバーライドや入力フィルターを指定できます。
|
||||
|
||||
### 基本的な使い方
|
||||
|
||||
シンプルなハンドオフの作成方法は以下の通りです。
|
||||
簡単なハンドオフは次のように作成できます。
|
||||
|
||||
```python
|
||||
from agents import Agent, handoff
|
||||
@@ -24,18 +30,22 @@ refund_agent = Agent(name="Refund agent")
|
||||
triage_agent = Agent(name="Triage agent", handoffs=[billing_agent, handoff(refund_agent)])
|
||||
```
|
||||
|
||||
1. エージェントを直接指定する(例:`billing_agent`)ことも、`handoff()` 関数を使うこともできます。
|
||||
1. エージェントを(`billing_agent` のように)直接使用することも、`handoff()` 関数を使用することもできます。
|
||||
|
||||
### `handoff()` 関数によるハンドオフのカスタマイズ
|
||||
|
||||
[`handoff()`][agents.handoffs.handoff] 関数では、さまざまなカスタマイズが可能です。
|
||||
[`handoff()`][agents.handoffs.handoff] 関数では、さまざまな要素をカスタマイズできます。
|
||||
|
||||
- `agent`: ハンドオフ先のエージェントです。
|
||||
- `tool_name_override`: デフォルトでは `Handoff.default_tool_name()` 関数が使われ、`transfer_to_<agent_name>` となります。これを上書きできます。
|
||||
- `tool_description_override`: デフォルトのツール説明(`Handoff.default_tool_description()`)を上書きできます。
|
||||
- `on_handoff`: ハンドオフが呼び出されたときに実行されるコールバック関数です。たとえば、ハンドオフが呼び出されたタイミングでデータ取得を開始するなどに便利です。この関数はエージェントコンテキストを受け取り、オプションで LLM が生成した入力も受け取れます。入力データは `input_type` パラメーターで制御します。
|
||||
- `input_type`: ハンドオフで期待される入力の型(オプション)です。
|
||||
- `input_filter`: 次のエージェントが受け取る入力をフィルタリングできます。詳細は下記をご覧ください。
|
||||
- `agent`: 処理のハンドオフ先となるエージェントです。
|
||||
- `tool_name_override`: 既定では `Handoff.default_tool_name()` 関数が使用され、これは `transfer_to_<agent_name>` に解決されます。これを上書きできます。
|
||||
- `tool_description_override`: `Handoff.default_tool_description()` の既定のツール説明を上書きします。
|
||||
- `on_handoff`: ハンドオフが呼び出されたときに実行されるコールバック関数です。ハンドオフが呼び出されることが分かった時点でデータ取得を開始する、といった用途に便利です。この関数はエージェントコンテキストを受け取り、任意で LLM が生成した入力も受け取れます。入力データは `input_type` パラメーターで制御されます。
|
||||
- `input_type`: ハンドオフのツール呼び出し引数のスキーマです。設定されている場合、解析済みのペイロードが `on_handoff` に渡されます。
|
||||
- `input_filter`: 次のエージェントが受け取る入力をフィルタリングできます。詳細は下記を参照してください。
|
||||
- `is_enabled`: ハンドオフが有効かどうかです。これはブール値、またはブール値を返す関数にできます。これにより、実行時にハンドオフを動的に有効化または無効化できます。
|
||||
- `nest_handoff_history`: RunConfig レベルの `nest_handoff_history` 設定に対する、呼び出しごとの任意のオーバーライドです。`None` の場合、アクティブな実行設定で定義された値が代わりに使用されます。
|
||||
|
||||
[`handoff()`][agents.handoffs.handoff] ヘルパーは、常に渡された特定の `agent` に制御を移します。複数の宛先候補がある場合は、宛先ごとに 1 つのハンドオフを登録し、モデルにそれらの中から選ばせてください。独自のハンドオフコードが、呼び出し時にどのエージェントを返すかを決定する必要がある場合にのみ、カスタム [`Handoff`][agents.handoffs.Handoff] を使用してください。
|
||||
|
||||
```python
|
||||
from agents import Agent, handoff, RunContextWrapper
|
||||
@@ -55,7 +65,7 @@ handoff_obj = handoff(
|
||||
|
||||
## ハンドオフ入力
|
||||
|
||||
状況によっては、LLM にハンドオフ時に何らかのデータを提供してほしい場合があります。たとえば、「エスカレーションエージェント」へのハンドオフを考えてみましょう。この場合、理由を提供して記録できるようにしたいことがあります。
|
||||
状況によっては、ハンドオフを呼び出す際に LLM に何らかのデータを提供させたい場合があります。たとえば、「エスカレーションエージェント」へのハンドオフを想像してください。理由を提供させて、それをログに記録したい場合があります。
|
||||
|
||||
```python
|
||||
from pydantic import BaseModel
|
||||
@@ -77,11 +87,44 @@ handoff_obj = handoff(
|
||||
)
|
||||
```
|
||||
|
||||
`input_type` は、ハンドオフツール呼び出し自体の引数を記述します。SDK はそのスキーマをハンドオフツールの `parameters` としてモデルに公開し、返された JSON をローカルで検証して、解析済みの値を `on_handoff` に渡します。
|
||||
|
||||
これは次のエージェントのメイン入力を置き換えるものではなく、別の宛先を選ぶものでもありません。[`handoff()`][agents.handoffs.handoff] ヘルパーは引き続き、ラップした特定のエージェントに転送し、受信側エージェントは [`input_filter`][agents.handoffs.Handoff.input_filter] またはネストされたハンドオフ履歴設定で変更しない限り、会話履歴を参照します。
|
||||
|
||||
`input_type` は [`RunContextWrapper.context`][agents.run_context.RunContextWrapper.context] とも別です。`input_type` は、ハンドオフ時にモデルが決定するメタデータに使用し、アプリケーション状態やローカルにすでにある依存関係には使用しないでください。
|
||||
|
||||
### `input_type` の使用場面
|
||||
|
||||
`input_type` は、ハンドオフで `reason`、`language`、`priority`、`summary` など、モデルが生成する小さなメタデータが必要な場合に使用します。たとえば、トリアージエージェントは `{ "reason": "duplicate_charge", "priority": "high" }` を付けて返金エージェントにハンドオフでき、`on_handoff` は返金エージェントが引き継ぐ前にそのメタデータをログに記録したり永続化したりできます。
|
||||
|
||||
目的が異なる場合は、別の仕組みを選択してください。
|
||||
|
||||
- 既存のアプリケーション状態と依存関係は [`RunContextWrapper.context`][agents.run_context.RunContextWrapper.context] に入れてください。[コンテキストガイド](context.md) を参照してください。
|
||||
- 受信側エージェントが参照する履歴を変更したい場合は、[`input_filter`][agents.handoffs.Handoff.input_filter]、[`RunConfig.nest_handoff_history`][agents.run.RunConfig.nest_handoff_history]、または [`RunConfig.handoff_history_mapper`][agents.run.RunConfig.handoff_history_mapper] を使用してください。
|
||||
- 複数の専門エージェント候補がある場合は、宛先ごとに 1 つのハンドオフを登録してください。`input_type` は選択されたハンドオフにメタデータを追加できますが、宛先間の振り分けは行いません。
|
||||
- 会話を転送せずにネストされた専門エージェントへ構造化入力を渡したい場合は、[`Agent.as_tool(parameters=...)`][agents.agent.Agent.as_tool] を優先してください。[ツール](tools.md#structured-input-for-tool-agents)を参照してください。
|
||||
|
||||
## 入力フィルター
|
||||
|
||||
ハンドオフが発生すると、新しいエージェントが会話を引き継ぎ、これまでの会話履歴全体を見ることができます。これを変更したい場合は、[`input_filter`][agents.handoffs.Handoff.input_filter] を設定できます。入力フィルターは、[`HandoffInputData`][agents.handoffs.HandoffInputData] を受け取り、新しい `HandoffInputData` を返す関数です。
|
||||
ハンドオフが発生すると、新しいエージェントが会話を引き継いだかのように、以前の会話履歴全体を参照できるようになります。これを変更したい場合は、[`input_filter`][agents.handoffs.Handoff.input_filter] を設定できます。入力フィルターは、[`HandoffInputData`][agents.handoffs.HandoffInputData] 経由で既存の入力を受け取り、新しい `HandoffInputData` を返す必要がある関数です。
|
||||
|
||||
よくあるパターン(たとえば履歴からすべてのツール呼び出しを削除するなど)は、[`agents.extensions.handoff_filters`][] で実装されています。
|
||||
[`HandoffInputData`][agents.handoffs.HandoffInputData] には次が含まれます。
|
||||
|
||||
- `input_history`: `Runner.run(...)` が開始される前の入力履歴です。
|
||||
- `pre_handoff_items`: ハンドオフが呼び出されたエージェントターンの前に生成された項目です。
|
||||
- `new_items`: 現在のターン中に生成された項目です。ハンドオフ呼び出しとハンドオフ出力項目を含みます。
|
||||
- `input_items`: `new_items` の代わりに次のエージェントへ転送する任意の項目です。`new_items` をセッション履歴用にそのまま保持しながら、モデル入力をフィルタリングできます。
|
||||
- `run_context`: ハンドオフが呼び出された時点でアクティブな [`RunContextWrapper`][agents.run_context.RunContextWrapper] です。
|
||||
|
||||
ネストされたハンドオフはオプトインのベータとして利用可能で、安定化中のため既定では無効です。[`RunConfig.nest_handoff_history`][agents.run.RunConfig.nest_handoff_history] を有効にすると、ランナーはそれまでのトランスクリプトを単一の assistant 要約メッセージにまとめ、それを `<CONVERSATION HISTORY>` ブロックでラップします。このブロックには、同じ実行中に複数のハンドオフが発生した場合、新しいターンが追加され続けます。完全な `input_filter` を書かずに生成されたメッセージを置き換えるには、[`RunConfig.handoff_history_mapper`][agents.run.RunConfig.handoff_history_mapper] 経由で独自のマッピング関数を提供できます。このオプトインは、ハンドオフと実行のどちらも明示的な `input_filter` を提供していない場合にのみ適用されるため、すでにペイロードをカスタマイズしている既存のコード(このリポジトリのコード例を含む)は、変更なしで現在の動作を維持します。単一のハンドオフについてネスト動作を上書きするには、[`handoff(...)`][agents.handoffs.handoff] に `nest_handoff_history=True` または `False` を渡します。これにより [`Handoff.nest_handoff_history`][agents.handoffs.Handoff.nest_handoff_history] が設定されます。生成された要約のラッパーテキストを変更するだけでよい場合は、エージェントを実行する前に [`set_conversation_history_wrappers`][agents.handoffs.set_conversation_history_wrappers](および必要に応じて [`reset_conversation_history_wrappers`][agents.handoffs.reset_conversation_history_wrappers])を呼び出してください。
|
||||
|
||||
ハンドオフとアクティブな [`RunConfig.handoff_input_filter`][agents.run.RunConfig.handoff_input_filter] の両方でフィルターが定義されている場合、その特定のハンドオフでは、ハンドオフごとの [`input_filter`][agents.handoffs.Handoff.input_filter] が優先されます。
|
||||
|
||||
!!! note
|
||||
|
||||
ハンドオフは単一の実行内に留まります。入力ガードレールは引き続きチェーン内の最初のエージェントにのみ適用され、出力ガードレールは最終出力を生成するエージェントにのみ適用されます。ワークフロー内の各カスタム関数ツール呼び出しの周辺でチェックが必要な場合は、ツールガードレールを使用してください。
|
||||
|
||||
一般的なパターン(たとえば履歴からすべてのツール呼び出しを削除するなど)は、[`agents.extensions.handoff_filters`][] に実装されています。
|
||||
|
||||
```python
|
||||
from agents import Agent, handoff
|
||||
@@ -95,11 +138,11 @@ handoff_obj = handoff(
|
||||
)
|
||||
```
|
||||
|
||||
1. これにより、`FAQ agent` が呼び出されたときに履歴からすべてのツールが自動的に削除されます。
|
||||
1. これにより、`FAQ agent` が呼び出されたときに、履歴からすべてのツールが自動的に削除されます。
|
||||
|
||||
## 推奨プロンプト
|
||||
|
||||
LLM がハンドオフを正しく理解できるようにするため、エージェントにハンドオフに関する情報を含めることを推奨します。[`agents.extensions.handoff_prompt.RECOMMENDED_PROMPT_PREFIX`][] で推奨されるプレフィックスを利用するか、[`agents.extensions.handoff_prompt.prompt_with_handoff_instructions`][] を呼び出して、推奨データをプロンプトに自動追加できます。
|
||||
LLM がハンドオフを適切に理解できるように、エージェントにハンドオフに関する情報を含めることをおすすめします。推奨プレフィックスは [`agents.extensions.handoff_prompt.RECOMMENDED_PROMPT_PREFIX`][] に用意されています。または、[`agents.extensions.handoff_prompt.prompt_with_handoff_instructions`][] を呼び出して、推奨データをプロンプトに自動的に追加できます。
|
||||
|
||||
```python
|
||||
from agents import Agent
|
||||
|
||||
@@ -0,0 +1,208 @@
|
||||
---
|
||||
search:
|
||||
exclude: true
|
||||
---
|
||||
# ヒューマンインザループ
|
||||
|
||||
human-in-the-loop (HITL) フローを使用すると、機密性の高いツール呼び出しが人によって承認または拒否されるまで、エージェントの実行を一時停止できます。ツールは承認が必要なタイミングを宣言し、実行結果は保留中の承認を中断として表面化し、`RunState` によって判断後の実行をシリアライズして再開できます。
|
||||
|
||||
この承認が表面化する範囲は実行全体であり、現在のトップレベルエージェントに限定されません。同じパターンは、ツールが現在のエージェントに属する場合、ハンドオフで到達したエージェントに属する場合、またはネストされた [`Agent.as_tool()`][agents.agent.Agent.as_tool] 実行に属する場合にも適用されます。ネストされた `Agent.as_tool()` の場合でも、中断は外側の実行に表面化するため、外側の `RunState` で承認または拒否し、元のトップレベル実行を再開します。
|
||||
|
||||
`Agent.as_tool()` では、承認は 2 つの異なる層で発生することがあります。エージェントツール自体が `Agent.as_tool(..., needs_approval=...)` によって承認を必要とする場合があり、ネストされたエージェント内のツールが、ネストされた実行の開始後に独自の承認を要求する場合もあります。どちらも同じ外側の実行の中断フローで処理されます。
|
||||
|
||||
このページでは、`interruptions` による手動承認フローに焦点を当てます。アプリがコード内で判断できる場合、一部のツールタイプはプログラムによる承認コールバックもサポートしているため、一時停止せずに実行を継続できます。
|
||||
|
||||
## 承認が必要なツールのマーク付け
|
||||
|
||||
常に承認を必須にするには `needs_approval` を `True` に設定し、呼び出しごとに判断するには async 関数を指定します。この呼び出し可能オブジェクトは、実行コンテキスト、解析済みのツールパラメーター、ツール呼び出し ID を受け取ります。
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner, function_tool
|
||||
|
||||
|
||||
@function_tool(needs_approval=True)
|
||||
async def cancel_order(order_id: int) -> str:
|
||||
return f"Cancelled order {order_id}"
|
||||
|
||||
|
||||
async def requires_review(_ctx, params, _call_id) -> bool:
|
||||
return "refund" in params.get("subject", "").lower()
|
||||
|
||||
|
||||
@function_tool(needs_approval=requires_review)
|
||||
async def send_email(subject: str, body: str) -> str:
|
||||
return f"Sent '{subject}'"
|
||||
|
||||
|
||||
agent = Agent(
|
||||
name="Support agent",
|
||||
instructions="Handle tickets and ask for approval when needed.",
|
||||
tools=[cancel_order, send_email],
|
||||
)
|
||||
```
|
||||
|
||||
`needs_approval` は [`function_tool`][agents.tool.function_tool]、[`Agent.as_tool`][agents.agent.Agent.as_tool]、[`ShellTool`][agents.tool.ShellTool]、[`ApplyPatchTool`][agents.tool.ApplyPatchTool] で利用できます。ローカル MCP サーバーも、[`MCPServerStdio`][agents.mcp.server.MCPServerStdio]、[`MCPServerSse`][agents.mcp.server.MCPServerSse]、[`MCPServerStreamableHttp`][agents.mcp.server.MCPServerStreamableHttp] の `require_approval` によって承認をサポートします。ホスト型 MCP サーバーは、`tool_config={"require_approval": "always"}` と任意の `on_approval_request` コールバックを指定した [`HostedMCPTool`][agents.tool.HostedMCPTool] によって承認をサポートします。シェルツールと apply_patch ツールは、中断を表面化させずに自動承認または自動拒否したい場合に、`on_approval` コールバックを受け付けます。
|
||||
|
||||
## 承認フローの仕組み
|
||||
|
||||
1. モデルがツール呼び出しを生成すると、ランナーはその承認ルール(`needs_approval`、`require_approval`、またはホスト型 MCP の同等設定)を評価します。
|
||||
2. そのツール呼び出しに対する承認判断がすでに [`RunContextWrapper`][agents.run_context.RunContextWrapper] に保存されている場合、ランナーはプロンプトを表示せずに続行します。呼び出しごとの承認は特定の呼び出し ID にスコープされます。実行の残りの期間におけるそのツールへの今後の呼び出しにも同じ判断を保持するには、`always_approve=True` または `always_reject=True` を渡します。
|
||||
3. それ以外の場合、実行は一時停止し、`RunResult.interruptions`(または `RunResultStreaming.interruptions`)に、`agent.name`、`tool_name`、`arguments` などの詳細を含む [`ToolApprovalItem`][agents.items.ToolApprovalItem] エントリが含まれます。これには、ハンドオフ後、またはネストされた `Agent.as_tool()` 実行内で要求された承認も含まれます。
|
||||
4. `result.to_state()` で実行結果を `RunState` に変換し、`state.approve(...)` または `state.reject(...)` を呼び出してから、`Runner.run(agent, state)` または `Runner.run_streamed(agent, state)` で再開します。ここで `agent` は、その実行の元のトップレベルエージェントです。
|
||||
5. 再開された実行は中断した箇所から続行し、新しい承認が必要になった場合はこのフローに再び入ります。
|
||||
|
||||
`always_approve=True` または `always_reject=True` で作成された固定的な判断は実行状態に保存されるため、後で同じ一時停止中の実行を再開する際に、`state.to_string()` / `RunState.from_string(...)` や `state.to_json()` / `RunState.from_json(...)` を経ても保持されます。
|
||||
|
||||
同じパスで保留中の承認をすべて解決する必要はありません。`interruptions` には、通常の関数ツール、ホスト型 MCP 承認、ネストされた `Agent.as_tool()` 承認が混在する場合があります。一部の項目だけを承認または拒否してから再実行すると、解決済みの呼び出しは続行できますが、未解決のものは `interruptions` に残り、実行を再び一時停止します。
|
||||
|
||||
## カスタム拒否メッセージ
|
||||
|
||||
デフォルトでは、拒否されたツール呼び出しは SDK の標準拒否テキストを実行内に返します。このメッセージは 2 つの層でカスタマイズできます。
|
||||
|
||||
- 実行全体のフォールバック: 実行全体にわたる承認拒否について、モデルから見えるデフォルトメッセージを制御するには、[`RunConfig.tool_error_formatter`][agents.run.RunConfig.tool_error_formatter] を設定します。
|
||||
- 呼び出しごとの上書き: 特定の拒否されたツール呼び出しで別のメッセージを表面化させたい場合は、`state.reject(...)` に `rejection_message=...` を渡します。
|
||||
|
||||
両方が指定された場合は、呼び出しごとの `rejection_message` が実行全体のフォーマッターより優先されます。
|
||||
|
||||
```python
|
||||
from agents import RunConfig, ToolErrorFormatterArgs
|
||||
|
||||
|
||||
def format_rejection(args: ToolErrorFormatterArgs[None]) -> str | None:
|
||||
if args.kind != "approval_rejected":
|
||||
return None
|
||||
return "Publish action was canceled because approval was rejected."
|
||||
|
||||
|
||||
run_config = RunConfig(tool_error_formatter=format_rejection)
|
||||
|
||||
# Later, while resolving a specific interruption:
|
||||
state.reject(
|
||||
interruption,
|
||||
rejection_message="Publish action was canceled because the reviewer denied approval.",
|
||||
)
|
||||
```
|
||||
|
||||
両方の層を組み合わせて示す完全な例については、[`examples/agent_patterns/human_in_the_loop_custom_rejection.py`](https://github.com/openai/openai-agents-python/tree/main/examples/agent_patterns/human_in_the_loop_custom_rejection.py) を参照してください。
|
||||
|
||||
## 自動承認判断
|
||||
|
||||
手動の `interruptions` は最も一般的なパターンですが、それだけではありません。
|
||||
|
||||
- ローカルの [`ShellTool`][agents.tool.ShellTool] と [`ApplyPatchTool`][agents.tool.ApplyPatchTool] は、`on_approval` を使用してコード内で即座に承認または拒否できます。
|
||||
- [`HostedMCPTool`][agents.tool.HostedMCPTool] は、`tool_config={"require_approval": "always"}` と `on_approval_request` を組み合わせて使用することで、同じ種類のプログラムによる判断を行えます。
|
||||
- 通常の [`function_tool`][agents.tool.function_tool] ツールと [`Agent.as_tool()`][agents.agent.Agent.as_tool] は、このページの手動中断フローを使用します。
|
||||
|
||||
これらのコールバックが判断を返すと、実行は人間の応答を待って一時停止せずに続行します。Realtime と音声セッション API については、[Realtime ガイド](realtime/guide.md)の承認フローを参照してください。
|
||||
|
||||
## ストリーミングとセッション
|
||||
|
||||
同じ中断フローはストリーミング実行でも機能します。ストリーミング実行が一時停止した後は、イテレーターが終了するまで [`RunResultStreaming.stream_events()`][agents.result.RunResultStreaming.stream_events] を消費し続け、[`RunResultStreaming.interruptions`][agents.result.RunResultStreaming.interruptions] を確認し、それらを解決して、再開後の出力もストリーミングし続けたい場合は [`Runner.run_streamed(...)`][agents.run.Runner.run_streamed] で再開します。このパターンのストリーミング版については、[ストリーミング](streaming.md)を参照してください。
|
||||
|
||||
セッションも使用している場合は、`RunState` から再開するときに同じセッションインスタンスを渡し続けるか、同じバックエンドストアを指す別のセッションオブジェクトを渡します。これにより、再開されたターンは同じ保存済み会話履歴に追加されます。セッションのライフサイクルの詳細については、[セッション](sessions/index.md)を参照してください。
|
||||
|
||||
## 例: 一時停止、承認、再開
|
||||
|
||||
以下のスニペットは JavaScript HITL ガイドに対応するものです。ツールが承認を必要とすると一時停止し、状態をディスクに永続化し、再読み込みして、判断を収集した後に再開します。
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
import json
|
||||
from pathlib import Path
|
||||
|
||||
from agents import Agent, Runner, RunState, function_tool
|
||||
|
||||
|
||||
async def needs_oakland_approval(_ctx, params, _call_id) -> bool:
|
||||
return "Oakland" in params.get("city", "")
|
||||
|
||||
|
||||
@function_tool(needs_approval=needs_oakland_approval)
|
||||
async def get_temperature(city: str) -> str:
|
||||
return f"The temperature in {city} is 20° Celsius"
|
||||
|
||||
|
||||
agent = Agent(
|
||||
name="Weather assistant",
|
||||
instructions="Answer weather questions with the provided tools.",
|
||||
tools=[get_temperature],
|
||||
)
|
||||
|
||||
STATE_PATH = Path(".cache/hitl_state.json")
|
||||
|
||||
|
||||
def prompt_approval(tool_name: str, arguments: str | None) -> bool:
|
||||
answer = input(f"Approve {tool_name} with {arguments}? [y/N]: ").strip().lower()
|
||||
return answer in {"y", "yes"}
|
||||
|
||||
|
||||
async def main() -> None:
|
||||
result = await Runner.run(agent, "What is the temperature in Oakland?")
|
||||
|
||||
while result.interruptions:
|
||||
# Persist the paused state.
|
||||
state = result.to_state()
|
||||
STATE_PATH.parent.mkdir(parents=True, exist_ok=True)
|
||||
STATE_PATH.write_text(state.to_string())
|
||||
|
||||
# Load the state later (could be a different process).
|
||||
stored = json.loads(STATE_PATH.read_text())
|
||||
state = await RunState.from_json(agent, stored)
|
||||
|
||||
for interruption in result.interruptions:
|
||||
approved = await asyncio.get_running_loop().run_in_executor(
|
||||
None, prompt_approval, interruption.name or "unknown_tool", interruption.arguments
|
||||
)
|
||||
if approved:
|
||||
state.approve(interruption, always_approve=False)
|
||||
else:
|
||||
state.reject(interruption)
|
||||
|
||||
result = await Runner.run(agent, state)
|
||||
|
||||
print(result.final_output)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
この例では、`prompt_approval` は `input()` を使用し、`run_in_executor(...)` で実行されるため同期的です。承認元がすでに非同期である場合(たとえば、HTTP リクエストや async データベースクエリ)、代わりに `async def` 関数を使用して直接 `await` できます。
|
||||
|
||||
承認待ちの間に出力をストリーミングするには、`Runner.run_streamed` を呼び出し、完了するまで `result.stream_events()` を消費してから、上記と同じ `result.to_state()` と再開手順に従います。
|
||||
|
||||
## リポジトリのパターンとコード例
|
||||
|
||||
- **ストリーミング承認**: `examples/agent_patterns/human_in_the_loop_stream.py` は、`stream_events()` を最後まで消費し、その後に保留中のツール呼び出しを承認してから `Runner.run_streamed(agent, state)` で再開する方法を示しています。
|
||||
- **カスタム拒否テキスト**: `examples/agent_patterns/human_in_the_loop_custom_rejection.py` は、承認が拒否されたときに、実行レベルの `tool_error_formatter` と呼び出しごとの `rejection_message` 上書きを組み合わせる方法を示しています。
|
||||
- **ツールとしてのエージェントの承認**: `Agent.as_tool(..., needs_approval=...)` は、委任されたエージェントタスクにレビューが必要な場合に、同じ中断フローを適用します。ネストされた中断も外側の実行に表面化するため、ネストされたエージェントではなく、元のトップレベルエージェントを再開します。
|
||||
- **ローカルシェルと apply_patch ツール**: `ShellTool` と `ApplyPatchTool` も `needs_approval` をサポートします。今後の呼び出しに対して判断をキャッシュするには、`state.approve(interruption, always_approve=True)` または `state.reject(..., always_reject=True)` を使用します。自動判断には `on_approval` を指定します(`examples/tools/shell.py` を参照)。手動判断には中断を処理します(`examples/tools/shell_human_in_the_loop.py` を参照)。ホスト型シェル環境は `needs_approval` または `on_approval` をサポートしていません。[ツールガイド](tools.md)を参照してください。
|
||||
- **ローカル MCP サーバー**: MCP ツール呼び出しを制御するには、`MCPServerStdio` / `MCPServerSse` / `MCPServerStreamableHttp` で `require_approval` を使用します(`examples/mcp/get_all_mcp_tools_example/main.py` と `examples/mcp/tool_filter_example/main.py` を参照)。
|
||||
- **ホスト型 MCP サーバー**: HITL を強制するには、`HostedMCPTool` で `require_approval` を `"always"` に設定し、必要に応じて自動承認または拒否のために `on_approval_request` を指定します(`examples/hosted_mcp/human_in_the_loop.py` と `examples/hosted_mcp/on_approval.py` を参照)。信頼できるサーバーには `"never"` を使用します(`examples/hosted_mcp/simple.py`)。
|
||||
- **セッションとメモリ**: 承認と会話履歴が複数ターンにわたって保持されるように、`Runner.run` にセッションを渡します。SQLite と OpenAI Conversations セッションのバリエーションは、`examples/memory/memory_session_hitl_example.py` と `examples/memory/openai_session_hitl_example.py` にあります。
|
||||
- **Realtime エージェント**: Realtime デモは、`RealtimeSession` 上の `approve_tool_call` / `reject_tool_call` を介してツール呼び出しを承認または拒否する WebSocket メッセージを公開しています(サーバー側ハンドラーについては `examples/realtime/app/server.py`、API サーフェスについては [Realtime ガイド](realtime/guide.md#tool-approvals)を参照)。
|
||||
|
||||
## 長時間にわたる承認
|
||||
|
||||
`RunState` は耐久性を持つように設計されています。保留中の作業をデータベースまたはキューに保存するには `state.to_json()` または `state.to_string()` を使用し、後で `RunState.from_json(...)` または `RunState.from_string(...)` で再作成します。
|
||||
|
||||
有用なシリアライズオプション:
|
||||
|
||||
- `context_serializer`: 非マッピングのコンテキストオブジェクトをシリアライズする方法をカスタマイズします。
|
||||
- `context_deserializer`: `RunState.from_json(...)` または `RunState.from_string(...)` で状態を読み込むときに、非マッピングのコンテキストオブジェクトを再構築します。
|
||||
- `strict_context=True`: コンテキストがすでに
|
||||
マッピングであるか、適切なシリアライザー/デシリアライザーを指定している場合を除き、シリアライズまたはデシリアライズを失敗させます。
|
||||
- `context_override`: 状態を読み込むときに、シリアライズされたコンテキストを置き換えます。これは、元のコンテキストオブジェクトを復元したくない場合に便利ですが、
|
||||
すでにシリアライズされたペイロードからそのコンテキストを削除するわけではありません。
|
||||
- `include_tracing_api_key=True`: 同じ認証情報でトレースのエクスポートを継続するために再開後の作業で必要な場合、
|
||||
シリアライズされたトレースペイロードにトレーシング API キーを含めます。
|
||||
|
||||
シリアライズされた実行状態には、アプリのコンテキストに加えて、承認、
|
||||
使用量、シリアライズされた `tool_input`、ネストされた agent-as-tool の再開、トレースメタデータ、サーバー管理の
|
||||
会話設定など、SDK が管理するランタイムメタデータが含まれます。シリアライズされた状態を保存または送信する予定がある場合は、
|
||||
`RunContextWrapper.context` を永続化データとして扱い、状態と一緒に移動することを意図している場合を除き、
|
||||
そこにシークレットを置かないでください。
|
||||
|
||||
## 保留中タスクのバージョニング
|
||||
|
||||
承認がしばらく保留状態のままになる可能性がある場合は、シリアライズされた状態と一緒に、エージェント定義または SDK のバージョンマーカーを保存してください。そうすれば、モデル、プロンプト、ツール定義が変更されたときの非互換性を避けるために、デシリアライズを対応するコードパスにルーティングできます。
|
||||
+68
-19
@@ -1,28 +1,54 @@
|
||||
---
|
||||
search:
|
||||
exclude: true
|
||||
---
|
||||
# OpenAI Agents SDK
|
||||
|
||||
[OpenAI Agents SDK](https://github.com/openai/openai-agents-python) は、非常に少ない抽象化で、軽量かつ使いやすいパッケージで エージェント 型 AI アプリを構築できる SDK です。これは、以前のエージェント向け実験プロジェクト [Swarm](https://github.com/openai/swarm/tree/main) の本番運用向けアップグレード版です。Agents SDK には、非常に少数の基本コンポーネントが含まれています。
|
||||
[OpenAI Agents SDK](https://github.com/openai/openai-agents-python) は、抽象化をほとんど持たない軽量で使いやすいパッケージで、エージェント型 AI アプリを構築できるようにします。これは、以前のエージェント向け実験プロジェクトである [Swarm](https://github.com/openai/swarm/tree/main) を本番環境対応に発展させたものです。Agents SDK は、非常に少数の基本コンポーネントで構成されています:
|
||||
|
||||
- **エージェント**: instructions と tools を備えた LLM
|
||||
- **ハンドオフ**:特定のタスクを他のエージェントに委任できる仕組み
|
||||
- **ガードレール**:エージェントへの入力を検証できる仕組み
|
||||
- **エージェント**: 指示とツールを備えた LLM です
|
||||
- **Agents as tools / ハンドオフ**: エージェントが特定のタスクを他のエージェントに委任できるようにします
|
||||
- **ガードレール**: エージェントの入力と出力の検証を可能にします
|
||||
|
||||
Python と組み合わせることで、これらの基本コンポーネントは tools と エージェント 間の複雑な関係を表現でき、急な学習コストなしに実用的なアプリケーションを構築できます。さらに、SDK には組み込みの **トレーシング** 機能があり、エージェント フローの可視化やデバッグ、評価、さらにはアプリケーション向けモデルのファインチューニングも可能です。
|
||||
Python と組み合わせることで、これらの基本コンポーネントは、ツールとエージェント間の複雑な関係を表現するのに十分強力であり、習得のハードルを高くすることなく実世界のアプリケーションを構築できます。さらに、SDK には組み込みの **トレーシング** が含まれており、エージェント型フローの可視化とデバッグ、評価、さらにはアプリケーション向けのモデルのファインチューニングも可能です。
|
||||
|
||||
## なぜ Agents SDK を使うのか
|
||||
## Agents SDK の利用理由
|
||||
|
||||
SDK には 2 つの設計原則があります。
|
||||
SDK の設計を支える原則は 2 つあります:
|
||||
|
||||
1. 使う価値のある十分な機能を持ちつつ、学習が容易なほど少ない基本コンポーネントで構成されていること。
|
||||
2. すぐに使い始められるが、動作を細かくカスタマイズできること。
|
||||
1. 使用する価値がある十分な機能を備えつつ、すばやく学べるだけの少数の基本コンポーネントに抑えること。
|
||||
2. そのままでも優れた動作をしつつ、何が起こるかを正確にカスタマイズできること。
|
||||
|
||||
SDK の主な特徴は以下の通りです。
|
||||
SDK の主な機能は次のとおりです:
|
||||
|
||||
- エージェントループ: tools の呼び出し、LLM への実行結果の送信、LLM が完了するまでのループ処理を自動で行います。
|
||||
- Python ファースト:新しい抽象化を学ぶ必要なく、Python の言語機能でエージェントのオーケストレーションや連携が可能です。
|
||||
- ハンドオフ:複数のエージェント間での調整や委任を実現する強力な機能です。
|
||||
- ガードレール:エージェントと並行して入力検証やチェックを実行し、チェックに失敗した場合は早期に処理を中断します。
|
||||
- 関数ツール:任意の Python 関数を自動でスキーマ生成・Pydantic ベースのバリデーション付きツールに変換できます。
|
||||
- トレーシング:組み込みのトレーシングでワークフローの可視化・デバッグ・モニタリングができ、OpenAI の評価・ファインチューニング・蒸留ツールも利用可能です。
|
||||
- **エージェントループ**: ツール呼び出しを処理し、結果を LLM に送り返し、タスクが完了するまで継続する組み込みのエージェントループです。
|
||||
- **Python ファースト**: 新しい抽象化を学ぶ必要なく、組み込みの言語機能を使ってエージェントをオーケストレーションし、連鎖させます。
|
||||
- **Agents as tools / ハンドオフ**: 複数のエージェント間で作業を調整し、委任するための強力な仕組みです。
|
||||
- **Sandbox エージェント**: マニフェストで定義されたファイル、Sandbox クライアントの選択、再開可能なサンドボックスセッションを備えた、実際の隔離ワークスペース内で専門エージェントを実行します。
|
||||
- **ガードレール**: エージェント実行と並行して入力検証と安全性チェックを実行し、チェックに通らない場合は即座に失敗として終了します。
|
||||
- **関数ツール**: スキーマの自動生成と Pydantic によるバリデーションにより、任意の Python 関数をツールに変換します。
|
||||
- **MCP サーバーのツール呼び出し**: 関数ツールと同じように動作する、組み込みの MCP サーバーツール統合です。
|
||||
- **セッション**: エージェントループ内で作業コンテキストを維持するための永続的なメモリレイヤーです。
|
||||
- **ヒューマンインザループ**: エージェント実行の各所に人間を関与させるための組み込みの仕組みです。
|
||||
- **トレーシング**: ワークフローを可視化、デバッグ、監視するための組み込みのトレーシングで、OpenAI の評価、ファインチューニング、蒸留ツール群をサポートします。
|
||||
- **Realtime エージェント**: `gpt-realtime-2` を使い、自動割り込み検出、コンテキスト管理、ガードレールなどを備えた強力な音声エージェントを構築します。
|
||||
|
||||
## Agents SDK と Responses API の選択
|
||||
|
||||
SDK は OpenAI モデルに対してデフォルトで Responses API を使用しますが、モデル呼び出しの周りに高レベルのランタイムを追加します。
|
||||
|
||||
次の場合は Responses API を直接使用します:
|
||||
|
||||
- ループ、ツールのディスパッチ、状態処理を自分で管理したい場合
|
||||
- ワークフローが短期間で、主にモデルの応答を返すことが目的の場合
|
||||
|
||||
次の場合は Agents SDK を使用します:
|
||||
|
||||
- ランタイムにターン、ツール実行、ガードレール、ハンドオフ、またはセッションを管理させたい場合
|
||||
- エージェントが成果物を生成する、または複数の協調したステップにわたって動作する必要がある場合
|
||||
- 実際のワークスペース、または [Sandbox エージェント](sandbox_agents.md) による再開可能な実行が必要な場合
|
||||
|
||||
アプリケーション全体でどちらか一方を選ぶ必要はありません。多くのアプリケーションでは、管理されたワークフローには SDK を使用し、低レベルの処理経路では Responses API を直接呼び出します。
|
||||
|
||||
## インストール
|
||||
|
||||
@@ -30,7 +56,7 @@ SDK の主な特徴は以下の通りです。
|
||||
pip install openai-agents
|
||||
```
|
||||
|
||||
## Hello World サンプル
|
||||
## Hello world の例
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner
|
||||
@@ -45,8 +71,31 @@ print(result.final_output)
|
||||
# Infinite loop's dance.
|
||||
```
|
||||
|
||||
(_実行する場合は、`OPENAI_API_KEY` 環境変数を設定してください_)
|
||||
(_これを実行する場合は、 `OPENAI_API_KEY` 環境変数を設定していることを確認してください_)
|
||||
|
||||
```bash
|
||||
export OPENAI_API_KEY=sk-...
|
||||
```
|
||||
```
|
||||
|
||||
## 開始ポイント
|
||||
|
||||
- [クイックスタート](quickstart.md) で、最初のテキストベースのエージェントを構築します。
|
||||
- 次に、[エージェントの実行](running_agents.md#choose-a-memory-strategy) で、ターン間で状態をどのように引き継ぐかを決定します。
|
||||
- タスクが実際のファイル、リポジトリ、またはエージェントごとに隔離されたワークスペース状態に依存する場合は、[Sandbox エージェントのクイックスタート](sandbox_agents.md) を参照してください。
|
||||
- ハンドオフとマネージャースタイルのオーケストレーションのどちらにするかを決める場合は、[エージェントオーケストレーション](multi_agent.md) を参照してください。
|
||||
|
||||
## パスの選択
|
||||
|
||||
実行したい作業は分かっているものの、どのページで説明されているか分からない場合は、この表を使用してください。
|
||||
|
||||
| 目的 | 参照先 |
|
||||
| --- | --- |
|
||||
| 最初のテキストエージェントを構築し、完全な 1 回の実行を確認する | [クイックスタート](quickstart.md) |
|
||||
| 関数ツール、OpenAI がホストするツール、または agents as tools を追加する | [ツール](tools.md) |
|
||||
| 実際の隔離ワークスペース内で、コーディング、レビュー、またはドキュメント処理のエージェントを実行する | [Sandbox エージェントのクイックスタート](sandbox_agents.md) and [Sandbox クライアント](sandbox/clients.md) |
|
||||
| ハンドオフとマネージャースタイルのオーケストレーションのどちらを使うか決める | [エージェントオーケストレーション](multi_agent.md) |
|
||||
| ターン間でメモリを保持する | [エージェントの実行](running_agents.md#choose-a-memory-strategy) and [セッション](sessions/index.md) |
|
||||
| OpenAI モデル、WebSocket トランスポート、または OpenAI 以外のプロバイダーを使用する | [モデル](models/index.md) |
|
||||
| 出力、実行アイテム、割り込み、再開状態を確認する | [実行結果](results.md) |
|
||||
| `gpt-realtime-2` を使って低レイテンシの音声エージェントを構築する | [Realtime エージェントのクイックスタート](realtime/quickstart.md) and [Realtime トランスポート](realtime/transport.md) |
|
||||
| 音声認識 / エージェント / 音声合成のパイプラインを構築する | [音声パイプラインのクイックスタート](voice/quickstart.md) |
|
||||
+456
-29
@@ -1,60 +1,487 @@
|
||||
---
|
||||
search:
|
||||
exclude: true
|
||||
---
|
||||
# Model context protocol (MCP)
|
||||
|
||||
[Model context protocol](https://modelcontextprotocol.io/introduction)(通称 MCP)は、 LLM にツールやコンテキストを提供するための方法です。MCP ドキュメントからの引用です:
|
||||
[Model context protocol](https://modelcontextprotocol.io/introduction) (MCP) は、アプリケーションが言語モデルにツールと
|
||||
コンテキストを公開する方法を標準化します。公式ドキュメントより:
|
||||
|
||||
> MCP は、アプリケーションが LLM にコンテキストを提供する方法を標準化するオープンプロトコルです。MCP を AI アプリケーションのための USB-C ポートのようなものと考えてください。USB-C がさまざまな周辺機器やアクセサリにデバイスを接続する標準的な方法を提供するのと同様に、MCP は AI モデルをさまざまなデータソースやツールに接続する標準的な方法を提供します。
|
||||
> MCP は、アプリケーションが LLM にコンテキストを提供する方法を標準化するオープンプロトコルです。MCP は AI
|
||||
> アプリケーションのための USB-C ポートのようなものだと考えてください。USB-C がデバイスをさまざまな周辺機器やアクセサリに接続する標準化された方法を提供するのと同じように、MCP
|
||||
> は AI モデルをさまざまなデータソースやツールに接続する標準化された方法を提供します。
|
||||
|
||||
Agents SDK は MCP をサポートしています。これにより、幅広い MCP サーバーを利用して、エージェントにツールを提供することができます。
|
||||
Agents Python SDK は複数の MCP トランスポートに対応しています。これにより、既存の MCP サーバーを再利用したり、独自に構築して
|
||||
ファイルシステム、HTTP、またはコネクターを基盤とするツールをエージェントに公開できます。
|
||||
|
||||
## MCP サーバー
|
||||
## MCP 統合の選択
|
||||
|
||||
現在、MCP 仕様では、使用するトランスポートメカニズムに基づいて 2 種類のサーバーが定義されています:
|
||||
MCP サーバーをエージェントに接続する前に、ツール呼び出しをどこで実行するか、どのトランスポートに到達できるかを決定してください。
|
||||
以下の表は、Python SDK がサポートする選択肢をまとめたものです。
|
||||
|
||||
1. **stdio** サーバーは、アプリケーションのサブプロセスとして実行されます。ローカルで実行されるものと考えることができます。
|
||||
2. **HTTP over SSE** サーバーはリモートで実行されます。URL を介して接続します。
|
||||
| 必要なこと | 推奨オプション |
|
||||
| ------------------------------------------------------------------------------------ | ----------------------------------------------------- |
|
||||
| OpenAI の Responses API に、モデルの代理で公開到達可能な MCP サーバーを呼び出させる| [`HostedMCPTool`][agents.tool.HostedMCPTool] 経由の **ホスト型 MCP サーバーツール** |
|
||||
| ローカルまたはリモートで実行する Streamable HTTP サーバーに接続する | [`MCPServerStreamableHttp`][agents.mcp.server.MCPServerStreamableHttp] 経由の **Streamable HTTP MCP サーバー** |
|
||||
| Server-Sent Events を用いた HTTP を実装しているサーバーと通信する | [`MCPServerSse`][agents.mcp.server.MCPServerSse] 経由の **SSE 付き HTTP MCP サーバー** |
|
||||
| ローカルプロセスを起動し、stdin/stdout 経由で通信する | [`MCPServerStdio`][agents.mcp.server.MCPServerStdio] 経由の **stdio MCP サーバー** |
|
||||
|
||||
これらのサーバーに接続するには、[`MCPServerStdio`][agents.mcp.server.MCPServerStdio] および [`MCPServerSse`][agents.mcp.server.MCPServerSse] クラスを使用できます。
|
||||
以下のセクションでは、各オプション、設定方法、あるトランスポートを別のトランスポートより優先すべき場合について説明します。
|
||||
|
||||
例えば、[公式 MCP ファイルシステムサーバー](https://www.npmjs.com/package/@modelcontextprotocol/server-filesystem) を使用する場合は、次のようになります。
|
||||
## エージェントレベルの MCP 設定
|
||||
|
||||
トランスポートの選択に加えて、`Agent.mcp_config` を設定することで、MCP ツールの準備方法を調整できます。
|
||||
|
||||
```python
|
||||
from agents import Agent
|
||||
|
||||
agent = Agent(
|
||||
name="Assistant",
|
||||
mcp_servers=[server],
|
||||
mcp_config={
|
||||
# Try to convert MCP tool schemas to strict JSON schema.
|
||||
"convert_schemas_to_strict": True,
|
||||
# If None, MCP tool failures are raised as exceptions instead of
|
||||
# returning model-visible error text.
|
||||
"failure_error_function": None,
|
||||
# Prefix local MCP tool names with their server name.
|
||||
"include_server_in_tool_names": True,
|
||||
},
|
||||
)
|
||||
```
|
||||
|
||||
注:
|
||||
|
||||
- `convert_schemas_to_strict` はベストエフォートです。スキーマを変換できない場合は、元のスキーマが使用されます。
|
||||
- `failure_error_function` は、MCP ツール呼び出しの失敗をモデルにどのように提示するかを制御します。
|
||||
- `failure_error_function` が未設定の場合、SDK はデフォルトのツールエラーフォーマッターを使用します。
|
||||
- サーバーレベルの `failure_error_function` は、そのサーバーについて `Agent.mcp_config["failure_error_function"]` を上書きします。
|
||||
- `include_server_in_tool_names` はオプトインです。有効にすると、各ローカル MCP ツールは、決定的なサーバープレフィックス付きの名前でモデルに公開されます。これにより、複数の MCP サーバーが同じ名前のツールを公開している場合の衝突を避けやすくなります。生成される名前は ASCII セーフで、関数ツール名の長さ制限内に収まり、同じエージェント上の既存のローカル関数ツール名と有効なハンドオフ名を避けます。SDK は引き続き、元のサーバー上で元の MCP ツール名を呼び出します。
|
||||
|
||||
## トランスポート間で共通するパターン
|
||||
|
||||
トランスポートを選択した後、多くの統合では同じような追加判断が必要です。
|
||||
|
||||
- ツールのサブセットのみを公開する方法([ツールフィルタリング](#tool-filtering))。
|
||||
- サーバーが再利用可能なプロンプトも提供するかどうか([プロンプト](#prompts))。
|
||||
- `list_tools()` をキャッシュすべきかどうか([キャッシュ](#caching))。
|
||||
- MCP アクティビティがトレースにどのように表示されるか([トレーシング](#tracing))。
|
||||
|
||||
ローカル MCP サーバー(`MCPServerStdio`、`MCPServerSse`、`MCPServerStreamableHttp`)では、承認ポリシーと呼び出しごとの `_meta` ペイロードも共通の概念です。Streamable HTTP セクションでは最も完全な例を示しており、同じパターンは他のローカルトランスポートにも適用できます。
|
||||
|
||||
## 1. ホスト型 MCP サーバーツール
|
||||
|
||||
ホスト型ツールでは、ツールの往復処理全体を OpenAI のインフラストラクチャに委ねます。コードでツールを一覧表示して呼び出す代わりに、
|
||||
[`HostedMCPTool`][agents.tool.HostedMCPTool] がサーバーラベル(および任意のコネクターメタデータ)を Responses API に転送します。
|
||||
モデルはリモートサーバーのツールを一覧表示し、Python プロセスへの追加のコールバックなしでそれらを呼び出します。ホスト型ツールは現在、
|
||||
Responses API のホスト型 MCP 統合をサポートする OpenAI モデルで動作します。
|
||||
|
||||
### 基本的なホスト型 MCP ツール
|
||||
|
||||
[`HostedMCPTool`][agents.tool.HostedMCPTool] をエージェントの `tools` リストに追加して、ホスト型ツールを作成します。`tool_config`
|
||||
dict は、REST API に送信する JSON と同じ構造です。
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
|
||||
from agents import Agent, HostedMCPTool, Runner
|
||||
|
||||
async def main() -> None:
|
||||
agent = Agent(
|
||||
name="Assistant",
|
||||
instructions="Use the DeepWiki hosted MCP server to inspect openai/openai-agents-python.",
|
||||
tools=[
|
||||
HostedMCPTool(
|
||||
tool_config={
|
||||
"type": "mcp",
|
||||
"server_label": "deepwiki",
|
||||
"server_url": "https://mcp.deepwiki.com/mcp",
|
||||
"require_approval": "never",
|
||||
}
|
||||
)
|
||||
],
|
||||
)
|
||||
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
"Which language is the repository openai/openai-agents-python written in?",
|
||||
)
|
||||
print(result.final_output)
|
||||
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
ホスト型サーバーはツールを自動的に公開します。`mcp_servers` に追加する必要はありません。
|
||||
|
||||
ホスト型ツール検索でホスト型 MCP サーバーを遅延読み込みしたい場合は、`tool_config["defer_loading"] = True` を設定し、[`ToolSearchTool`][agents.tool.ToolSearchTool] をエージェントに追加します。これは OpenAI Responses モデルでのみサポートされます。ツール検索の完全な設定と制約については、[ツール](tools.md#hosted-tool-search)を参照してください。
|
||||
|
||||
### ホスト型 MCP 実行結果のストリーミング
|
||||
|
||||
ホスト型ツールは、関数ツールとまったく同じ方法で実行結果のストリーミングをサポートします。`Runner.run_streamed` を使用して、
|
||||
モデルがまだ動作している間に、増分 MCP 出力を受け取ります。
|
||||
|
||||
```python
|
||||
result = Runner.run_streamed(agent, "Summarise this repository's top languages")
|
||||
async for event in result.stream_events():
|
||||
if event.type == "run_item_stream_event":
|
||||
print(f"Received: {event.item}")
|
||||
print(result.final_output)
|
||||
```
|
||||
|
||||
### 任意の承認フロー
|
||||
|
||||
サーバーが機密性の高い操作を実行できる場合は、各ツール実行の前に人間による承認またはプログラムによる承認を要求できます。
|
||||
`tool_config` の `require_approval` に、単一のポリシー(`"always"`、`"never"`)またはツール名からポリシーへの
|
||||
マッピング dict を設定します。Python 内で判断するには、`on_approval_request` コールバックを提供します。
|
||||
|
||||
```python
|
||||
from agents import MCPToolApprovalFunctionResult, MCPToolApprovalRequest
|
||||
|
||||
SAFE_TOOLS = {"read_wiki_structure", "read_wiki_contents", "ask_question"}
|
||||
|
||||
def approve_tool(request: MCPToolApprovalRequest) -> MCPToolApprovalFunctionResult:
|
||||
if request.data.name in SAFE_TOOLS:
|
||||
return {"approve": True}
|
||||
return {"approve": False, "reason": "Escalate to a human reviewer"}
|
||||
|
||||
agent = Agent(
|
||||
name="Assistant",
|
||||
tools=[
|
||||
HostedMCPTool(
|
||||
tool_config={
|
||||
"type": "mcp",
|
||||
"server_label": "deepwiki",
|
||||
"server_url": "https://mcp.deepwiki.com/mcp",
|
||||
"require_approval": "always",
|
||||
},
|
||||
on_approval_request=approve_tool,
|
||||
)
|
||||
],
|
||||
)
|
||||
```
|
||||
|
||||
コールバックは同期または非同期にでき、モデルが実行を続けるために承認データを必要とするたびに呼び出されます。
|
||||
|
||||
### コネクター対応のホスト型サーバー
|
||||
|
||||
ホスト型 MCP は OpenAI コネクターもサポートしています。`server_url` を指定する代わりに、`connector_id` とアクセストークンを指定します。
|
||||
Responses API が認証を処理し、ホスト型サーバーがコネクターのツールを公開します。
|
||||
|
||||
```python
|
||||
import os
|
||||
|
||||
HostedMCPTool(
|
||||
tool_config={
|
||||
"type": "mcp",
|
||||
"server_label": "google_calendar",
|
||||
"connector_id": "connector_googlecalendar",
|
||||
"authorization": os.environ["GOOGLE_CALENDAR_AUTHORIZATION"],
|
||||
"require_approval": "never",
|
||||
}
|
||||
)
|
||||
```
|
||||
|
||||
ストリーミング、承認、コネクターを含む、完全に動作するホスト型ツールサンプルは、
|
||||
[`examples/hosted_mcp`](https://github.com/openai/openai-agents-python/tree/main/examples/hosted_mcp) にあります。
|
||||
|
||||
## 2. Streamable HTTP MCP サーバー
|
||||
|
||||
ネットワーク接続を自分で管理したい場合は、
|
||||
[`MCPServerStreamableHttp`][agents.mcp.server.MCPServerStreamableHttp] を使用します。Streamable HTTP サーバーは、トランスポートを自分で制御する場合や、
|
||||
レイテンシーを低く保ちながら自分のインフラストラクチャ内でサーバーを実行したい場合に最適です。
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
import os
|
||||
|
||||
from agents import Agent, Runner
|
||||
from agents.mcp import MCPServerStreamableHttp
|
||||
from agents.model_settings import ModelSettings
|
||||
|
||||
async def main() -> None:
|
||||
token = os.environ["MCP_SERVER_TOKEN"]
|
||||
async with MCPServerStreamableHttp(
|
||||
name="Streamable HTTP Python Server",
|
||||
params={
|
||||
"url": "http://localhost:8000/mcp",
|
||||
"headers": {"Authorization": f"Bearer {token}"},
|
||||
"timeout": 10,
|
||||
},
|
||||
cache_tools_list=True,
|
||||
max_retry_attempts=3,
|
||||
) as server:
|
||||
agent = Agent(
|
||||
name="Assistant",
|
||||
instructions="Use the MCP tools to answer the questions.",
|
||||
mcp_servers=[server],
|
||||
model_settings=ModelSettings(tool_choice="required"),
|
||||
)
|
||||
|
||||
result = await Runner.run(agent, "Add 7 and 22.")
|
||||
print(result.final_output)
|
||||
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
コンストラクターは追加オプションを受け付けます。
|
||||
|
||||
- `client_session_timeout_seconds` は HTTP 読み取りタイムアウトを制御します。
|
||||
- `use_structured_content` は、テキスト出力より `tool_result.structured_content` を優先するかどうかを切り替えます。
|
||||
- `max_retry_attempts` と `retry_backoff_seconds_base` は、`list_tools()` と `call_tool()` に自動再試行を追加します。
|
||||
- `tool_filter` により、ツールのサブセットのみを公開できます([ツールフィルタリング](#tool-filtering)を参照)。
|
||||
- `require_approval` は、ローカル MCP ツールで人間参加型の承認ポリシーを有効にします。
|
||||
- `failure_error_function` は、モデルに見える MCP ツール失敗メッセージをカスタマイズします。代わりにエラーを発生させるには、`None` に設定します。
|
||||
- `tool_meta_resolver` は、`call_tool()` の前に呼び出しごとの MCP `_meta` ペイロードを注入します。
|
||||
|
||||
### ローカル MCP サーバーの承認ポリシー
|
||||
|
||||
`MCPServerStdio`、`MCPServerSse`、`MCPServerStreamableHttp` はすべて `require_approval` を受け付けます。
|
||||
|
||||
サポートされる形式:
|
||||
|
||||
- すべてのツールに対する `"always"` または `"never"`。
|
||||
- `True` / `False`(always/never と同等)。
|
||||
- ツールごとのマップ。例: `{"delete_file": "always", "read_file": "never"}`。
|
||||
- グループ化されたオブジェクト:
|
||||
`{"always": {"tool_names": [...]}, "never": {"tool_names": [...]}}`。
|
||||
|
||||
```python
|
||||
async with MCPServerStreamableHttp(
|
||||
name="Filesystem MCP",
|
||||
params={"url": "http://localhost:8000/mcp"},
|
||||
require_approval={"always": {"tool_names": ["delete_file"]}},
|
||||
) as server:
|
||||
...
|
||||
```
|
||||
|
||||
完全な一時停止/再開フローについては、[人間参加型](human_in_the_loop.md)および `examples/mcp/get_all_mcp_tools_example/main.py` を参照してください。
|
||||
|
||||
### `tool_meta_resolver` による呼び出しごとのメタデータ
|
||||
|
||||
MCP サーバーが `_meta` 内にリクエストメタデータ(たとえば、テナント ID やトレースコンテキスト)を想定している場合は、`tool_meta_resolver` を使用します。以下の例では、`Runner.run(...)` に `context` として `dict` を渡すことを想定しています。
|
||||
|
||||
```python
|
||||
from agents.mcp import MCPServerStreamableHttp, MCPToolMetaContext
|
||||
|
||||
|
||||
def resolve_meta(context: MCPToolMetaContext) -> dict[str, str] | None:
|
||||
run_context_data = context.run_context.context or {}
|
||||
tenant_id = run_context_data.get("tenant_id")
|
||||
if tenant_id is None:
|
||||
return None
|
||||
return {"tenant_id": str(tenant_id), "source": "agents-sdk"}
|
||||
|
||||
|
||||
server = MCPServerStreamableHttp(
|
||||
name="Metadata-aware MCP",
|
||||
params={"url": "http://localhost:8000/mcp"},
|
||||
tool_meta_resolver=resolve_meta,
|
||||
)
|
||||
```
|
||||
|
||||
実行コンテキストが Pydantic モデル、dataclass、またはカスタムクラスの場合は、属性アクセスでテナント ID を読み取ってください。
|
||||
|
||||
### MCP ツール出力: テキストと画像
|
||||
|
||||
MCP ツールが画像コンテンツを返すと、SDK はそれを画像ツール出力エントリーに自動的にマッピングします。テキスト/画像が混在するレスポンスは出力項目のリストとして転送されるため、エージェントは通常の関数ツールからの画像出力を利用するのと同じ方法で MCP 画像実行結果を利用できます。
|
||||
|
||||
## 3. SSE 付き HTTP MCP サーバー
|
||||
|
||||
!!! warning
|
||||
|
||||
MCP プロジェクトでは Server-Sent Events トランスポートが非推奨になりました。新しい統合には Streamable HTTP または stdio を優先し、SSE はレガシーサーバー向けにのみ維持してください。
|
||||
|
||||
MCP サーバーが SSE 付き HTTP トランスポートを実装している場合は、
|
||||
[`MCPServerSse`][agents.mcp.server.MCPServerSse] をインスタンス化します。トランスポートを除けば、API は Streamable HTTP サーバーと同一です。
|
||||
|
||||
```python
|
||||
|
||||
from agents import Agent, Runner
|
||||
from agents.model_settings import ModelSettings
|
||||
from agents.mcp import MCPServerSse
|
||||
|
||||
workspace_id = "demo-workspace"
|
||||
|
||||
async with MCPServerSse(
|
||||
name="SSE Python Server",
|
||||
params={
|
||||
"url": "http://localhost:8000/sse",
|
||||
"headers": {"X-Workspace": workspace_id},
|
||||
},
|
||||
cache_tools_list=True,
|
||||
) as server:
|
||||
agent = Agent(
|
||||
name="Assistant",
|
||||
mcp_servers=[server],
|
||||
model_settings=ModelSettings(tool_choice="required"),
|
||||
)
|
||||
result = await Runner.run(agent, "What's the weather in Tokyo?")
|
||||
print(result.final_output)
|
||||
```
|
||||
|
||||
## 4. stdio MCP サーバー
|
||||
|
||||
ローカルサブプロセスとして実行される MCP サーバーには、[`MCPServerStdio`][agents.mcp.server.MCPServerStdio] を使用します。SDK は
|
||||
プロセスを起動し、パイプを開いたままにし、コンテキストマネージャーを抜けると自動的に閉じます。このオプションは、素早い概念実証や、
|
||||
サーバーがコマンドラインエントリーポイントのみを公開している場合に役立ちます。
|
||||
|
||||
```python
|
||||
from pathlib import Path
|
||||
from agents import Agent, Runner
|
||||
from agents.mcp import MCPServerStdio
|
||||
|
||||
current_dir = Path(__file__).parent
|
||||
samples_dir = current_dir / "sample_files"
|
||||
|
||||
async with MCPServerStdio(
|
||||
name="Filesystem Server via npx",
|
||||
params={
|
||||
"command": "npx",
|
||||
"args": ["-y", "@modelcontextprotocol/server-filesystem", str(samples_dir)],
|
||||
},
|
||||
) as server:
|
||||
agent = Agent(
|
||||
name="Assistant",
|
||||
instructions="Use the files in the sample directory to answer questions.",
|
||||
mcp_servers=[server],
|
||||
)
|
||||
result = await Runner.run(agent, "List the files available to you.")
|
||||
print(result.final_output)
|
||||
```
|
||||
|
||||
## 5. MCP サーバーマネージャー
|
||||
|
||||
複数の MCP サーバーがある場合は、`MCPServerManager` を使用して事前に接続し、接続済みのサブセットをエージェントに公開します。
|
||||
コンストラクターオプションと再接続の動作については、[MCPServerManager API リファレンス](ref/mcp/manager.md)を参照してください。
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner
|
||||
from agents.mcp import MCPServerManager, MCPServerStreamableHttp
|
||||
|
||||
servers = [
|
||||
MCPServerStreamableHttp(name="calendar", params={"url": "http://localhost:8000/mcp"}),
|
||||
MCPServerStreamableHttp(name="docs", params={"url": "http://localhost:8001/mcp"}),
|
||||
]
|
||||
|
||||
async with MCPServerManager(servers) as manager:
|
||||
agent = Agent(
|
||||
name="Assistant",
|
||||
instructions="Use MCP tools when they help.",
|
||||
mcp_servers=manager.active_servers,
|
||||
)
|
||||
result = await Runner.run(agent, "Which MCP tools are available?")
|
||||
print(result.final_output)
|
||||
```
|
||||
|
||||
主な動作:
|
||||
|
||||
- `active_servers` には、`drop_failed_servers=True`(デフォルト)の場合、正常に接続されたサーバーのみが含まれます。
|
||||
- 失敗は `failed_servers` と `errors` で追跡されます。
|
||||
- 最初の接続失敗で例外を発生させるには、`strict=True` を設定します。
|
||||
- 失敗したサーバーを再試行するには `reconnect(failed_only=True)` を呼び出し、すべてのサーバーを再起動するには `reconnect(failed_only=False)` を呼び出します。
|
||||
- ライフサイクルの動作を調整するには、`connect_timeout_seconds`、`cleanup_timeout_seconds`、`connect_in_parallel` を使用します。
|
||||
|
||||
## 共通のサーバー機能
|
||||
|
||||
以下のセクションは、MCP サーバートランスポート全体に適用されます(正確な API サーフェスはサーバークラスによって異なります)。
|
||||
|
||||
## ツールフィルタリング
|
||||
|
||||
各 MCP サーバーはツールフィルターをサポートしているため、エージェントに必要な関数だけを公開できます。フィルタリングは、
|
||||
構築時にも実行ごとに動的にも行えます。
|
||||
|
||||
### 静的ツールフィルタリング
|
||||
|
||||
単純な許可/ブロックリストを設定するには、[`create_static_tool_filter`][agents.mcp.create_static_tool_filter] を使用します。
|
||||
|
||||
```python
|
||||
from pathlib import Path
|
||||
|
||||
from agents.mcp import MCPServerStdio, create_static_tool_filter
|
||||
|
||||
samples_dir = Path("/path/to/files")
|
||||
|
||||
filesystem_server = MCPServerStdio(
|
||||
params={
|
||||
"command": "npx",
|
||||
"args": ["-y", "@modelcontextprotocol/server-filesystem", str(samples_dir)],
|
||||
},
|
||||
tool_filter=create_static_tool_filter(allowed_tool_names=["read_file", "write_file"]),
|
||||
)
|
||||
```
|
||||
|
||||
`allowed_tool_names` と `blocked_tool_names` の両方が指定された場合、SDK はまず許可リストを適用し、その後、残ったセットから
|
||||
ブロックされたツールを削除します。
|
||||
|
||||
### 動的ツールフィルタリング
|
||||
|
||||
より複雑なロジックには、[`ToolFilterContext`][agents.mcp.ToolFilterContext] を受け取る呼び出し可能オブジェクトを渡します。呼び出し可能オブジェクトは
|
||||
同期または非同期にでき、ツールを公開すべき場合に `True` を返します。
|
||||
|
||||
```python
|
||||
from pathlib import Path
|
||||
|
||||
from agents.mcp import MCPServerStdio, ToolFilterContext
|
||||
|
||||
samples_dir = Path("/path/to/files")
|
||||
|
||||
async def context_aware_filter(context: ToolFilterContext, tool) -> bool:
|
||||
if context.agent.name == "Code Reviewer" and tool.name.startswith("danger_"):
|
||||
return False
|
||||
return True
|
||||
|
||||
async with MCPServerStdio(
|
||||
params={
|
||||
"command": "npx",
|
||||
"args": ["-y", "@modelcontextprotocol/server-filesystem", samples_dir],
|
||||
}
|
||||
"args": ["-y", "@modelcontextprotocol/server-filesystem", str(samples_dir)],
|
||||
},
|
||||
tool_filter=context_aware_filter,
|
||||
) as server:
|
||||
tools = await server.list_tools()
|
||||
...
|
||||
```
|
||||
|
||||
## MCP サーバーの利用
|
||||
フィルターコンテキストは、アクティブな `run_context`、ツールを要求している `agent`、および `server_name` を公開します。
|
||||
|
||||
MCP サーバーはエージェントに追加できます。Agents SDK は、エージェントが実行されるたびに MCP サーバーの `list_tools()` を呼び出します。これにより、LLM は MCP サーバーのツールを認識できるようになります。LLM が MCP サーバーのツールを呼び出すと、SDK はそのサーバーの `call_tool()` を呼び出します。
|
||||
## プロンプト
|
||||
|
||||
MCP サーバーは、エージェントの指示を動的に生成するプロンプトも提供できます。プロンプトをサポートするサーバーは、次の 2 つの
|
||||
メソッドを公開します。
|
||||
|
||||
- `list_prompts()` は、利用可能なプロンプトテンプレートを列挙します。
|
||||
- `get_prompt(name, arguments)` は、必要に応じてパラメーター付きで具体的なプロンプトを取得します。
|
||||
|
||||
```python
|
||||
from agents import Agent
|
||||
|
||||
agent=Agent(
|
||||
name="Assistant",
|
||||
instructions="Use the tools to achieve the task",
|
||||
mcp_servers=[mcp_server_1, mcp_server_2]
|
||||
prompt_result = await server.get_prompt(
|
||||
"generate_code_review_instructions",
|
||||
{"focus": "security vulnerabilities", "language": "python"},
|
||||
)
|
||||
instructions = prompt_result.messages[0].content.text
|
||||
|
||||
agent = Agent(
|
||||
name="Code Reviewer",
|
||||
instructions=instructions,
|
||||
mcp_servers=[server],
|
||||
)
|
||||
```
|
||||
|
||||
## キャッシュ
|
||||
|
||||
エージェントが実行されるたびに、MCP サーバーの `list_tools()` が呼び出されます。特にサーバーがリモートの場合、これはレイテンシの原因となることがあります。ツールリストを自動的にキャッシュするには、[`MCPServerStdio`][agents.mcp.server.MCPServerStdio] および [`MCPServerSse`][agents.mcp.server.MCPServerSse] の両方に `cache_tools_list=True` を渡すことができます。ツールリストが変更されないことが確実な場合のみ、この設定を行ってください。
|
||||
|
||||
キャッシュを無効化したい場合は、サーバーで `invalidate_tools_cache()` を呼び出すことができます。
|
||||
|
||||
## エンドツーエンドの code examples
|
||||
|
||||
[examples/mcp](https://github.com/openai/openai-agents-python/tree/main/examples/mcp) で、完全な動作 code examples をご覧いただけます。
|
||||
エージェントの各実行では、各 MCP サーバーで `list_tools()` が呼び出されます。リモートサーバーでは顕著なレイテンシーが発生する可能性があるため、すべての MCP
|
||||
サーバークラスは `cache_tools_list` オプションを公開しています。ツール定義が頻繁に変わらないと確信できる場合にのみ、`True` に設定してください。後で最新のリストを強制的に取得するには、サーバーインスタンスで `invalidate_tools_cache()` を呼び出します。
|
||||
|
||||
## トレーシング
|
||||
|
||||
[トレーシング](../tracing.md) は、以下を含む MCP の操作を自動的に記録します:
|
||||
[トレーシング](./tracing.md)は、以下を含む MCP アクティビティを自動的にキャプチャします。
|
||||
|
||||
1. MCP サーバーへのツールリスト取得の呼び出し
|
||||
2. 関数呼び出しに関する MCP 関連情報
|
||||
1. ツールを一覧表示するための MCP サーバーへの呼び出し。
|
||||
2. ツール呼び出しに関する MCP 関連情報。
|
||||
|
||||

|
||||

|
||||
|
||||
## 参考資料
|
||||
|
||||
- [Model Context Protocol](https://modelcontextprotocol.io/) – 仕様と設計ガイド。
|
||||
- [examples/mcp](https://github.com/openai/openai-agents-python/tree/main/examples/mcp) – 実行可能な stdio、SSE、Streamable HTTP サンプル。
|
||||
- [examples/hosted_mcp](https://github.com/openai/openai-agents-python/tree/main/examples/hosted_mcp) – 承認とコネクターを含む完全なホスト型 MCP デモ。
|
||||
@@ -1,106 +0,0 @@
|
||||
# モデル
|
||||
|
||||
Agents SDK には、OpenAI モデルの 2 種類のサポートが標準で用意されています。
|
||||
|
||||
- **推奨**: [`OpenAIResponsesModel`][agents.models.openai_responses.OpenAIResponsesModel] は、新しい [Responses API](https://platform.openai.com/docs/api-reference/responses) を使って OpenAI API を呼び出します。
|
||||
- [`OpenAIChatCompletionsModel`][agents.models.openai_chatcompletions.OpenAIChatCompletionsModel] は、[Chat Completions API](https://platform.openai.com/docs/api-reference/chat) を使って OpenAI API を呼び出します。
|
||||
|
||||
## モデルの組み合わせ
|
||||
|
||||
1 つのワークフロー内で、各エージェントごとに異なるモデルを使いたい場合があります。たとえば、トリアージには小型で高速なモデルを使い、複雑なタスクにはより大きく高性能なモデルを使うことができます。[`Agent`][agents.Agent] を設定する際、以下のいずれかの方法で特定のモデルを選択できます。
|
||||
|
||||
1. OpenAI モデル名を直接渡す。
|
||||
2. 任意のモデル名と、その名前を Model インスタンスにマッピングできる [`ModelProvider`][agents.models.interface.ModelProvider] を渡す。
|
||||
3. [`Model`][agents.models.interface.Model] 実装を直接指定する。
|
||||
|
||||
!!!note
|
||||
|
||||
SDK は [`OpenAIResponsesModel`][agents.models.openai_responses.OpenAIResponsesModel] と [`OpenAIChatCompletionsModel`][agents.models.openai_chatcompletions.OpenAIChatCompletionsModel] の両方の形状をサポートしていますが、各ワークフローで 1 つのモデル形状のみを使うことを推奨します。なぜなら、2 つの形状はサポートする機能やツールが異なるためです。ワークフローでモデル形状を組み合わせて使う場合は、利用するすべての機能が両方で利用可能かご確認ください。
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner, AsyncOpenAI, OpenAIChatCompletionsModel
|
||||
import asyncio
|
||||
|
||||
spanish_agent = Agent(
|
||||
name="Spanish agent",
|
||||
instructions="You only speak Spanish.",
|
||||
model="o3-mini", # (1)!
|
||||
)
|
||||
|
||||
english_agent = Agent(
|
||||
name="English agent",
|
||||
instructions="You only speak English",
|
||||
model=OpenAIChatCompletionsModel( # (2)!
|
||||
model="gpt-4o",
|
||||
openai_client=AsyncOpenAI()
|
||||
),
|
||||
)
|
||||
|
||||
triage_agent = Agent(
|
||||
name="Triage agent",
|
||||
instructions="Handoff to the appropriate agent based on the language of the request.",
|
||||
handoffs=[spanish_agent, english_agent],
|
||||
model="gpt-3.5-turbo",
|
||||
)
|
||||
|
||||
async def main():
|
||||
result = await Runner.run(triage_agent, input="Hola, ¿cómo estás?")
|
||||
print(result.final_output)
|
||||
```
|
||||
|
||||
1. OpenAI モデル名を直接設定します。
|
||||
2. [`Model`][agents.models.interface.Model] 実装を指定します。
|
||||
|
||||
エージェントで使用するモデルをさらに細かく設定したい場合は、[`ModelSettings`][agents.models.interface.ModelSettings] を渡すことができます。これにより、temperature などのオプションのモデル設定パラメーターを指定できます。
|
||||
|
||||
```python
|
||||
from agents import Agent, ModelSettings
|
||||
|
||||
english_agent = Agent(
|
||||
name="English agent",
|
||||
instructions="You only speak English",
|
||||
model="gpt-4o",
|
||||
model_settings=ModelSettings(temperature=0.1),
|
||||
)
|
||||
```
|
||||
|
||||
## 他の LLM プロバイダーの利用
|
||||
|
||||
他の LLM プロバイダーは、3 つの方法で利用できます([こちら](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/) に code examples があります)。
|
||||
|
||||
1. [`set_default_openai_client`][agents.set_default_openai_client] は、`AsyncOpenAI` のインスタンスを LLM クライアントとしてグローバルに利用したい場合に便利です。これは、LLM プロバイダーが OpenAI 互換の API エンドポイントを持ち、`base_url` と `api_key` を設定できる場合に使います。[examples/model_providers/custom_example_global.py](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/custom_example_global.py) に設定例があります。
|
||||
2. [`ModelProvider`][agents.models.interface.ModelProvider] は `Runner.run` レベルで利用します。これにより、「この実行のすべてのエージェントでカスタムモデルプロバイダーを使う」と指定できます。[examples/model_providers/custom_example_provider.py](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/custom_example_provider.py) に設定例があります。
|
||||
3. [`Agent.model`][agents.agent.Agent.model] で、特定のエージェントインスタンスにモデルを指定できます。これにより、エージェントごとに異なるプロバイダーを組み合わせて使うことができます。[examples/model_providers/custom_example_agent.py](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/custom_example_agent.py) に設定例があります。
|
||||
|
||||
`platform.openai.com` の API キーがない場合は、`set_tracing_disabled()` でトレーシングを無効にするか、[別のトレーシングプロセッサー](tracing.md) を設定することを推奨します。
|
||||
|
||||
!!! note
|
||||
|
||||
これらの code examples では Chat Completions API/モデルを使っています。なぜなら、ほとんどの LLM プロバイダーはまだ Responses API をサポートしていないためです。もし LLM プロバイダーが Responses API をサポートしている場合は、Responses の利用を推奨します。
|
||||
|
||||
## 他の LLM プロバイダー利用時のよくある問題
|
||||
|
||||
### Tracing クライアントの 401 エラー
|
||||
|
||||
トレーシングに関連するエラーが発生した場合、これはトレースが OpenAI サーバーにアップロードされるため、OpenAI API キーがないことが原因です。解決方法は 3 つあります。
|
||||
|
||||
1. トレーシングを完全に無効化する: [`set_tracing_disabled(True)`][agents.set_tracing_disabled]。
|
||||
2. トレーシング用の OpenAI キーを設定する: [`set_tracing_export_api_key(...)`][agents.set_tracing_export_api_key]。この API キーはトレースのアップロードのみに使われ、[platform.openai.com](https://platform.openai.com/) のものが必要です。
|
||||
3. OpenAI 以外のトレースプロセッサーを使う。[トレーシングのドキュメント](tracing.md#custom-tracing-processors) をご覧ください。
|
||||
|
||||
### Responses API サポート
|
||||
|
||||
SDK はデフォルトで Responses API を使いますが、ほとんどの他の LLM プロバイダーはまだ対応していません。そのため、404 エラーなどが発生する場合があります。解決方法は 2 つあります。
|
||||
|
||||
1. [`set_default_openai_api("chat_completions")`][agents.set_default_openai_api] を呼び出します。これは、環境変数で `OPENAI_API_KEY` と `OPENAI_BASE_URL` を設定している場合に有効です。
|
||||
2. [`OpenAIChatCompletionsModel`][agents.models.openai_chatcompletions.OpenAIChatCompletionsModel] を使います。[こちら](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/) に code examples があります。
|
||||
|
||||
### structured outputs サポート
|
||||
|
||||
一部のモデルプロバイダーは [structured outputs](https://platform.openai.com/docs/guides/structured-outputs) をサポートしていません。その場合、次のようなエラーが発生することがあります。
|
||||
|
||||
```
|
||||
BadRequestError: Error code: 400 - {'error': {'message': "'response_format.type' : value is not one of the allowed values ['text','json_object']", 'type': 'invalid_request_error'}}
|
||||
```
|
||||
|
||||
これは一部のモデルプロバイダーの制限で、JSON 出力には対応していても、出力に使う `json_schema` を指定できない場合があります。現在この問題の修正に取り組んでいますが、JSON schema 出力をサポートしているプロバイダーの利用を推奨します。そうでない場合、不正な JSON によりアプリが頻繁に動作しなくなる可能性があります。
|
||||
@@ -0,0 +1,514 @@
|
||||
---
|
||||
search:
|
||||
exclude: true
|
||||
---
|
||||
# モデル
|
||||
|
||||
Agents SDK には、OpenAI モデルの標準サポートが 2 種類用意されています。
|
||||
|
||||
- **推奨**: 新しい [Responses API](https://platform.openai.com/docs/api-reference/responses) を使用して OpenAI API を呼び出す [`OpenAIResponsesModel`][agents.models.openai_responses.OpenAIResponsesModel]。
|
||||
- [Chat Completions API](https://platform.openai.com/docs/api-reference/chat) を使用して OpenAI API を呼び出す [`OpenAIChatCompletionsModel`][agents.models.openai_chatcompletions.OpenAIChatCompletionsModel]。
|
||||
|
||||
## モデル設定の選択
|
||||
|
||||
セットアップに合う最もシンプルな方法から始めてください。
|
||||
|
||||
| 実現したいこと | 推奨される方法 | 詳細 |
|
||||
| --- | --- | --- |
|
||||
| OpenAI モデルのみを使用する | デフォルトの OpenAI プロバイダーを Responses モデルパスで使用する | [OpenAI モデル](#openai-models) |
|
||||
| websocket トランスポート経由で OpenAI Responses API を使用する | Responses モデルパスを維持し、websocket トランスポートを有効にする | [Responses WebSocket トランスポート](#responses-websocket-transport) |
|
||||
| 1 つの非 OpenAI プロバイダーを使用する | 組み込みのプロバイダー統合ポイントから始める | [非 OpenAI モデル](#non-openai-models) |
|
||||
| エージェント間でモデルまたはプロバイダーを混在させる | 実行ごと、またはエージェントごとにプロバイダーを選択し、機能差を確認する | [1 つのワークフローでのモデルの混在](#mixing-models-in-one-workflow) と [プロバイダー間でのモデルの混在](#mixing-models-across-providers) |
|
||||
| 高度な OpenAI Responses リクエスト設定を調整する | OpenAI Responses パスで `ModelSettings` を使用する | [高度な OpenAI Responses 設定](#advanced-openai-responses-settings) |
|
||||
| 非 OpenAI または混在プロバイダーのルーティングにサードパーティアダプターを使用する | サポートされているベータ版アダプターを比較し、リリース予定のプロバイダーパスを検証する | [サードパーティアダプター](#third-party-adapters) |
|
||||
|
||||
## OpenAI モデル
|
||||
|
||||
ほとんどの OpenAI のみのアプリでは、デフォルトの OpenAI プロバイダーで文字列のモデル名を使用し、Responses モデルパスを維持する方法が推奨されます。
|
||||
|
||||
`Agent` の初期化時にモデルを指定しない場合、デフォルトモデルが使用されます。低レイテンシのエージェントワークフロー向けに、現在のデフォルトは [`gpt-5.4-mini`](https://developers.openai.com/api/docs/models/gpt-5.4-mini) で、`reasoning.effort="none"` と `verbosity="low"` が設定されています。アクセス権がある場合は、明示的な `model_settings` を維持しつつ、より高品質な [`gpt-5.5`](https://developers.openai.com/api/docs/models/gpt-5.5) をエージェントに設定することを推奨します。
|
||||
|
||||
[`gpt-5.5`](https://developers.openai.com/api/docs/models/gpt-5.5) などの他のモデルに切り替えたい場合、エージェントを設定する方法は 2 つあります。
|
||||
|
||||
### デフォルトモデル
|
||||
|
||||
まず、カスタムモデルを設定していないすべてのエージェントで特定のモデルを一貫して使用したい場合は、エージェントを実行する前に `OPENAI_DEFAULT_MODEL` 環境変数を設定します。
|
||||
|
||||
```bash
|
||||
export OPENAI_DEFAULT_MODEL=gpt-5.5
|
||||
python3 my_awesome_agent.py
|
||||
```
|
||||
|
||||
次に、`RunConfig` を通じて実行のデフォルトモデルを設定できます。エージェントにモデルを設定しない場合、この実行のモデルが使用されます。
|
||||
|
||||
```python
|
||||
from agents import Agent, RunConfig, Runner
|
||||
|
||||
agent = Agent(
|
||||
name="Assistant",
|
||||
instructions="You're a helpful agent.",
|
||||
)
|
||||
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
"Hello",
|
||||
run_config=RunConfig(model="gpt-5.5"),
|
||||
)
|
||||
```
|
||||
|
||||
#### GPT-5 モデル
|
||||
|
||||
この方法で [`gpt-5.5`](https://developers.openai.com/api/docs/models/gpt-5.5) などの GPT-5 モデルを使用すると、SDK はデフォルトの `ModelSettings` を適用します。ほとんどのユースケースで最も適切に動作する設定が適用されます。デフォルトモデルの reasoning effort を調整するには、独自の `ModelSettings` を渡します。
|
||||
|
||||
```python
|
||||
from openai.types.shared import Reasoning
|
||||
from agents import Agent, ModelSettings
|
||||
|
||||
my_agent = Agent(
|
||||
name="My Agent",
|
||||
instructions="You're a helpful agent.",
|
||||
# If OPENAI_DEFAULT_MODEL=gpt-5.5 is set, passing only model_settings works.
|
||||
# It's also fine to pass a GPT-5 model name explicitly:
|
||||
model="gpt-5.5",
|
||||
model_settings=ModelSettings(reasoning=Reasoning(effort="high"), verbosity="low")
|
||||
)
|
||||
```
|
||||
|
||||
低レイテンシにするには、GPT-5 モデルで `reasoning.effort="none"` を使用することを推奨します。
|
||||
|
||||
#### ComputerTool のモデル選択
|
||||
|
||||
エージェントに [`ComputerTool`][agents.tool.ComputerTool] が含まれる場合、実際の Responses リクエストで有効なモデルによって、SDK が送信する computer-tool ペイロードが決まります。明示的な `gpt-5.5` リクエストでは GA 組み込み `computer` ツールが使用され、明示的な `computer-use-preview` リクエストでは従来の `computer_use_preview` ペイロードが維持されます。
|
||||
|
||||
主な例外は、プロンプト管理の呼び出しです。プロンプトテンプレートがモデルを保持し、SDK がリクエストから `model` を省略する場合、SDK は、プロンプトが固定しているモデルを推測しないように、プレビュー互換の computer ペイロードをデフォルトにします。そのフローで GA パスを維持するには、リクエストで `model="gpt-5.5"` を明示するか、`ModelSettings(tool_choice="computer")` または `ModelSettings(tool_choice="computer_use")` で GA セレクターを強制します。
|
||||
|
||||
登録済みの [`ComputerTool`][agents.tool.ComputerTool] がある場合、`tool_choice="computer"`、`"computer_use"`、`"computer_use_preview"` は、有効なリクエストモデルに一致する組み込みセレクターに正規化されます。`ComputerTool` が登録されていない場合、これらの文字列は通常の関数名と同様に動作し続けます。
|
||||
|
||||
プレビュー互換のリクエストでは、`environment` と表示寸法を事前にシリアライズする必要があります。そのため、[`ComputerProvider`][agents.tool.ComputerProvider] ファクトリーを使用するプロンプト管理フローでは、具体的な `Computer` または `AsyncComputer` インスタンスを渡すか、リクエスト送信前に GA セレクターを強制する必要があります。移行の詳細については、[ツール](../tools.md#computertool-and-the-responses-computer-tool)を参照してください。
|
||||
|
||||
#### 非 GPT-5 モデル
|
||||
|
||||
カスタムの `model_settings` なしで非 GPT-5 モデル名を渡すと、SDK は任意のモデルと互換性のある汎用の `ModelSettings` に戻します。
|
||||
|
||||
### Responses のみのツール検索機能
|
||||
|
||||
次のツール機能は、OpenAI Responses モデルでのみサポートされています。
|
||||
|
||||
- [`ToolSearchTool`][agents.tool.ToolSearchTool]
|
||||
- [`tool_namespace()`][agents.tool.tool_namespace]
|
||||
- `@function_tool(defer_loading=True)` とその他の遅延読み込み Responses ツールサーフェス
|
||||
|
||||
これらの機能は、Chat Completions モデルおよび非 Responses バックエンドでは拒否されます。遅延読み込みツールを使用する場合は、エージェントに `ToolSearchTool()` を追加し、裸の名前空間名や遅延専用の関数名を強制するのではなく、`auto` または `required` のツール選択を通じてモデルにツールを読み込ませてください。セットアップの詳細と現在の制約については、[ツール](../tools.md#hosted-tool-search)を参照してください。
|
||||
|
||||
### Responses WebSocket トランスポート
|
||||
|
||||
デフォルトでは、OpenAI Responses API リクエストは HTTP トランスポートを使用します。OpenAI ベースのモデルを使用する場合は、websocket トランスポートを有効にできます。
|
||||
|
||||
#### 基本設定
|
||||
|
||||
```python
|
||||
from agents import set_default_openai_responses_transport
|
||||
|
||||
set_default_openai_responses_transport("websocket")
|
||||
```
|
||||
|
||||
これは、デフォルトの OpenAI プロバイダーによって解決される OpenAI Responses モデル(`"gpt-5.5"` などの文字列モデル名を含む)に影響します。
|
||||
|
||||
トランスポートの選択は、SDK がモデル名をモデルインスタンスに解決するときに行われます。具体的な [`Model`][agents.models.interface.Model] オブジェクトを渡す場合、そのトランスポートはすでに固定されています。[`OpenAIResponsesWSModel`][agents.models.openai_responses.OpenAIResponsesWSModel] は websocket を使用し、[`OpenAIResponsesModel`][agents.models.openai_responses.OpenAIResponsesModel] は HTTP を使用し、[`OpenAIChatCompletionsModel`][agents.models.openai_chatcompletions.OpenAIChatCompletionsModel] は Chat Completions のままです。`RunConfig(model_provider=...)` を渡す場合、グローバルデフォルトではなく、そのプロバイダーがトランスポート選択を制御します。
|
||||
|
||||
#### プロバイダーまたは実行レベルの設定
|
||||
|
||||
プロバイダーごと、または実行ごとに websocket トランスポートを設定することもできます。
|
||||
|
||||
```python
|
||||
from agents import Agent, OpenAIProvider, RunConfig, Runner
|
||||
|
||||
provider = OpenAIProvider(
|
||||
use_responses_websocket=True,
|
||||
# Optional; if omitted, OPENAI_WEBSOCKET_BASE_URL is used when set.
|
||||
websocket_base_url="wss://your-proxy.example/v1",
|
||||
# Optional low-level websocket keepalive settings.
|
||||
responses_websocket_options={"ping_interval": 20.0, "ping_timeout": 60.0},
|
||||
)
|
||||
|
||||
agent = Agent(name="Assistant")
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
"Hello",
|
||||
run_config=RunConfig(model_provider=provider),
|
||||
)
|
||||
```
|
||||
|
||||
OpenAI ベースのプロバイダーは、任意のエージェント登録設定も受け付けます。これは、OpenAI の設定で harness ID などのプロバイダーレベルの登録メタデータが想定されている場合の高度なオプションです。
|
||||
|
||||
```python
|
||||
from agents import (
|
||||
Agent,
|
||||
OpenAIAgentRegistrationConfig,
|
||||
OpenAIProvider,
|
||||
RunConfig,
|
||||
Runner,
|
||||
)
|
||||
|
||||
provider = OpenAIProvider(
|
||||
use_responses_websocket=True,
|
||||
agent_registration=OpenAIAgentRegistrationConfig(harness_id="your-harness-id"),
|
||||
)
|
||||
|
||||
agent = Agent(name="Assistant")
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
"Hello",
|
||||
run_config=RunConfig(model_provider=provider),
|
||||
)
|
||||
```
|
||||
|
||||
#### `MultiProvider` による高度なルーティング
|
||||
|
||||
プレフィックスベースのモデルルーティングが必要な場合(たとえば 1 つの実行で `openai/...` と `any-llm/...` のモデル名を混在させる場合)は、[`MultiProvider`][agents.MultiProvider] を使用し、そこで `openai_use_responses_websocket=True` を設定します。
|
||||
|
||||
`MultiProvider` は、歴史的なデフォルトを 2 つ保持しています。
|
||||
|
||||
- `openai/...` は OpenAI プロバイダーのエイリアスとして扱われるため、`openai/gpt-4.1` はモデル `gpt-4.1` としてルーティングされます。
|
||||
- 不明なプレフィックスは、パススルーされるのではなく `UserError` を発生させます。
|
||||
|
||||
OpenAI プロバイダーを、リテラルな名前空間付きモデル ID を想定する OpenAI 互換エンドポイントに向ける場合は、パススルー動作を明示的に有効にします。websocket が有効なセットアップでは、`MultiProvider` でも `openai_use_responses_websocket=True` を維持してください。
|
||||
|
||||
```python
|
||||
from agents import Agent, MultiProvider, RunConfig, Runner
|
||||
|
||||
provider = MultiProvider(
|
||||
openai_base_url="https://openrouter.ai/api/v1",
|
||||
openai_api_key="...",
|
||||
openai_use_responses_websocket=True,
|
||||
openai_prefix_mode="model_id",
|
||||
unknown_prefix_mode="model_id",
|
||||
)
|
||||
|
||||
agent = Agent(
|
||||
name="Assistant",
|
||||
instructions="Be concise.",
|
||||
model="openai/gpt-4.1",
|
||||
)
|
||||
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
"Hello",
|
||||
run_config=RunConfig(model_provider=provider),
|
||||
)
|
||||
```
|
||||
|
||||
バックエンドがリテラルな `openai/...` 文字列を想定する場合は、`openai_prefix_mode="model_id"` を使用します。バックエンドが `openrouter/openai/gpt-4.1-mini` など、他の名前空間付きモデル ID を想定する場合は、`unknown_prefix_mode="model_id"` を使用します。これらのオプションは、websocket トランスポート外の `MultiProvider` でも機能します。この例では、このセクションで説明しているトランスポート設定の一部であるため、websocket を有効なままにしています。同じオプションは [`responses_websocket_session()`][agents.responses_websocket_session] でも利用できます。
|
||||
|
||||
`MultiProvider` 経由でルーティングしながら同じプロバイダーレベルの登録メタデータが必要な場合は、`openai_agent_registration=OpenAIAgentRegistrationConfig(...)` を渡すと、基盤となる OpenAI プロバイダーに転送されます。
|
||||
|
||||
カスタムの OpenAI 互換エンドポイントまたはプロキシを使用する場合、websocket トランスポートには互換性のある websocket `/responses` エンドポイントも必要です。そのようなセットアップでは、`websocket_base_url` を明示的に設定する必要がある場合があります。
|
||||
|
||||
#### 注記
|
||||
|
||||
- これは websocket トランスポート上の Responses API であり、[Realtime API](../realtime/guide.md) ではありません。Responses websocket `/responses` エンドポイントをサポートしていない限り、Chat Completions や非 OpenAI プロバイダーには適用されません。
|
||||
- 環境でまだ利用できない場合は、`websockets` パッケージをインストールしてください。
|
||||
- websocket トランスポートを有効にした後、[`Runner.run_streamed()`][agents.run.Runner.run_streamed] を直接使用できます。複数ターンのワークフローで、ターン間(および入れ子の agent-as-tool 呼び出し)で同じ websocket 接続を再利用したい場合は、[`responses_websocket_session()`][agents.responses_websocket_session] ヘルパーを推奨します。[エージェントの実行](../running_agents.md)ガイドと [`examples/basic/stream_ws.py`](https://github.com/openai/openai-agents-python/tree/main/examples/basic/stream_ws.py) を参照してください。
|
||||
- 長い推論ターンやレイテンシの急増があるネットワークでは、`responses_websocket_options` で websocket keepalive の動作をカスタマイズしてください。遅延した pong フレームを許容するには `ping_timeout` を増やすか、ping を有効にしたままハートビートタイムアウトを無効にするには `ping_timeout=None` を設定します。websocket レイテンシより信頼性が重要な場合は、HTTP/SSE トランスポートを優先してください。
|
||||
|
||||
## 非 OpenAI モデル
|
||||
|
||||
非 OpenAI プロバイダーが必要な場合は、SDK の組み込みプロバイダー統合ポイントから始めてください。多くのセットアップでは、サードパーティアダプターを追加しなくてもこれで十分です。各パターンの例は [examples/model_providers](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/) にあります。
|
||||
|
||||
### 非 OpenAI プロバイダーの統合方法
|
||||
|
||||
| アプローチ | 使用する場合 | 範囲 |
|
||||
| --- | --- | --- |
|
||||
| [`set_default_openai_client`][agents.set_default_openai_client] | 1 つの OpenAI 互換エンドポイントを、ほとんどまたはすべてのエージェントのデフォルトにしたい場合 | グローバルデフォルト |
|
||||
| [`ModelProvider`][agents.models.interface.ModelProvider] | 1 つのカスタムプロバイダーを 1 回の実行に適用したい場合 | 実行ごと |
|
||||
| [`Agent.model`][agents.agent.Agent.model] | 異なるエージェントに異なるプロバイダーや具体的なモデルオブジェクトが必要な場合 | エージェントごと |
|
||||
| サードパーティアダプター | 組み込みパスでは提供されない、アダプター管理のプロバイダー対応範囲やルーティングが必要な場合 | [サードパーティアダプター](#third-party-adapters)を参照 |
|
||||
|
||||
これらの組み込みパスで、他の LLM プロバイダーを統合できます。
|
||||
|
||||
1. [`set_default_openai_client`][agents.set_default_openai_client] は、`AsyncOpenAI` のインスタンスを LLM クライアントとしてグローバルに使用したい場合に便利です。これは、LLM プロバイダーに OpenAI 互換 API エンドポイントがあり、`base_url` と `api_key` を設定できる場合に使用します。設定可能な例については、[examples/model_providers/custom_example_global.py](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/custom_example_global.py) を参照してください。
|
||||
2. [`ModelProvider`][agents.models.interface.ModelProvider] は `Runner.run` レベルです。これにより、「この実行内のすべてのエージェントにカスタムモデルプロバイダーを使用する」と指定できます。設定可能な例については、[examples/model_providers/custom_example_provider.py](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/custom_example_provider.py) を参照してください。
|
||||
3. [`Agent.model`][agents.agent.Agent.model] により、特定の Agent インスタンスでモデルを指定できます。これにより、異なるエージェントに対して異なるプロバイダーを自由に組み合わせられます。設定可能な例については、[examples/model_providers/custom_example_agent.py](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/custom_example_agent.py) を参照してください。
|
||||
|
||||
`platform.openai.com` からの API キーがない場合は、`set_tracing_disabled()` でトレーシングを無効にするか、[別のトレーシングプロセッサー](../tracing.md)を設定することを推奨します。
|
||||
|
||||
``` python
|
||||
from agents import Agent, AsyncOpenAI, OpenAIChatCompletionsModel, set_tracing_disabled
|
||||
|
||||
set_tracing_disabled(disabled=True)
|
||||
|
||||
client = AsyncOpenAI(api_key="Api_Key", base_url="Base URL of Provider")
|
||||
model = OpenAIChatCompletionsModel(model="Model_Name", openai_client=client)
|
||||
|
||||
agent= Agent(name="Helping Agent", instructions="You are a Helping Agent", model=model)
|
||||
```
|
||||
|
||||
!!! note
|
||||
|
||||
これらの例では、Chat Completions API/モデルを使用しています。多くの LLM プロバイダーは、まだ Responses API をサポートしていないためです。LLM プロバイダーが Responses API をサポートしている場合は、Responses の使用を推奨します。
|
||||
|
||||
## 1 つのワークフローでのモデルの混在
|
||||
|
||||
単一のワークフロー内で、エージェントごとに異なるモデルを使用したい場合があります。たとえば、トリアージには小さく高速なモデルを使用し、複雑なタスクにはより大きく高性能なモデルを使用できます。[`Agent`][agents.Agent] を設定する際、次のいずれかの方法で特定のモデルを選択できます。
|
||||
|
||||
1. モデル名を渡す。
|
||||
2. 任意のモデル名と、その名前を Model インスタンスにマッピングできる [`ModelProvider`][agents.models.interface.ModelProvider] を渡す。
|
||||
3. [`Model`][agents.models.interface.Model] 実装を直接提供する。
|
||||
|
||||
!!! note
|
||||
|
||||
SDK は [`OpenAIResponsesModel`][agents.models.openai_responses.OpenAIResponsesModel] と [`OpenAIChatCompletionsModel`][agents.models.openai_chatcompletions.OpenAIChatCompletionsModel] の両方の形状をサポートしていますが、2 つの形状でサポートされる機能とツールのセットが異なるため、各ワークフローでは単一のモデル形状を使用することを推奨します。ワークフローでモデル形状を組み合わせる必要がある場合は、使用するすべての機能が両方で利用できることを確認してください。
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner, AsyncOpenAI, OpenAIChatCompletionsModel
|
||||
import asyncio
|
||||
|
||||
spanish_agent = Agent(
|
||||
name="Spanish agent",
|
||||
instructions="You only speak Spanish.",
|
||||
model="gpt-5-mini", # (1)!
|
||||
)
|
||||
|
||||
english_agent = Agent(
|
||||
name="English agent",
|
||||
instructions="You only speak English",
|
||||
model=OpenAIChatCompletionsModel( # (2)!
|
||||
model="gpt-5-nano",
|
||||
openai_client=AsyncOpenAI()
|
||||
),
|
||||
)
|
||||
|
||||
triage_agent = Agent(
|
||||
name="Triage agent",
|
||||
instructions="Handoff to the appropriate agent based on the language of the request.",
|
||||
handoffs=[spanish_agent, english_agent],
|
||||
model="gpt-5.5",
|
||||
)
|
||||
|
||||
async def main():
|
||||
result = await Runner.run(triage_agent, input="Hola, ¿cómo estás?")
|
||||
print(result.final_output)
|
||||
```
|
||||
|
||||
1. OpenAI モデルの名前を直接設定します。
|
||||
2. [`Model`][agents.models.interface.Model] 実装を提供します。
|
||||
|
||||
エージェントで使用するモデルをさらに設定したい場合は、temperature などの任意のモデル設定パラメーターを提供する [`ModelSettings`][agents.models.interface.ModelSettings] を渡せます。
|
||||
|
||||
```python
|
||||
from agents import Agent, ModelSettings
|
||||
|
||||
english_agent = Agent(
|
||||
name="English agent",
|
||||
instructions="You only speak English",
|
||||
model="gpt-4.1",
|
||||
model_settings=ModelSettings(temperature=0.1),
|
||||
)
|
||||
```
|
||||
|
||||
## 高度な OpenAI Responses 設定
|
||||
|
||||
OpenAI Responses パスを使用していて、より詳細な制御が必要な場合は、`ModelSettings` から始めてください。
|
||||
|
||||
### 一般的な高度な `ModelSettings` オプション
|
||||
|
||||
OpenAI Responses API を使用している場合、いくつかのリクエストフィールドにはすでに直接対応する `ModelSettings` フィールドがあるため、それらに `extra_args` は不要です。
|
||||
|
||||
- `parallel_tool_calls`: 同じターン内の複数のツール呼び出しを許可または禁止します。
|
||||
- `truncation`: コンテキストがあふれる場合に失敗するのではなく、Responses API が最も古い会話アイテムを削除できるようにするには、`"auto"` を設定します。
|
||||
- `store`: 生成されたレスポンスを後で取得できるようにサーバー側に保存するかどうかを制御します。これは、レスポンス ID に依存する後続ワークフローや、`store=False` のときにローカル入力へのフォールバックが必要になる可能性があるセッション圧縮フローで重要です。
|
||||
- `context_management`: `compact_threshold` を使用した Responses 圧縮など、サーバー側のコンテキスト処理を設定します。
|
||||
- `prompt_cache_retention`: たとえば `"24h"` を使って、キャッシュされたプロンプトプレフィックスをより長く保持します。
|
||||
- `response_include`: `web_search_call.action.sources`、`file_search_call.results`、`reasoning.encrypted_content` など、よりリッチなレスポンスペイロードをリクエストします。
|
||||
- `top_logprobs`: 出力テキストの top-token logprobs をリクエストします。SDK は `message.output_text.logprobs` も自動的に追加します。
|
||||
- `retry`: モデル呼び出しに対する runner 管理の再試行設定を有効にします。[Runner 管理の再試行](#runner-managed-retries)を参照してください。
|
||||
|
||||
```python
|
||||
from agents import Agent, ModelSettings
|
||||
|
||||
research_agent = Agent(
|
||||
name="Research agent",
|
||||
model="gpt-5.5",
|
||||
model_settings=ModelSettings(
|
||||
parallel_tool_calls=False,
|
||||
truncation="auto",
|
||||
store=True,
|
||||
context_management=[{"type": "compaction", "compact_threshold": 200000}],
|
||||
prompt_cache_retention="24h",
|
||||
response_include=["web_search_call.action.sources"],
|
||||
top_logprobs=5,
|
||||
),
|
||||
)
|
||||
```
|
||||
|
||||
`store=False` を設定すると、Responses API はそのレスポンスを後でサーバー側から取得できるようには保持しません。これはステートレス、またはゼロデータ保持スタイルのフローに便利ですが、通常ならレスポンス ID を再利用する機能が、代わりにローカル管理の状態に依存する必要があることも意味します。たとえば、[`OpenAIResponsesCompactionSession`][agents.memory.openai_responses_compaction_session.OpenAIResponsesCompactionSession] は、最後のレスポンスが保存されていない場合、デフォルトの `"auto"` 圧縮パスを入力ベースの圧縮に切り替えます。[Sessions ガイド](../sessions/index.md#openai-responses-compaction-sessions)を参照してください。
|
||||
|
||||
サーバー側圧縮は、[`OpenAIResponsesCompactionSession`][agents.memory.openai_responses_compaction_session.OpenAIResponsesCompactionSession] とは異なります。`context_management=[{"type": "compaction", "compact_threshold": ...}]` は各 Responses API リクエストと一緒に送信され、レンダリングされたコンテキストがしきい値を超えると、API はレスポンスの一部として圧縮アイテムを出力できます。`OpenAIResponsesCompactionSession` はターン間でスタンドアロンの `responses.compact` エンドポイントを呼び出し、ローカルセッション履歴を書き換えます。
|
||||
|
||||
### `extra_args` の受け渡し
|
||||
|
||||
プロバイダー固有、または SDK がまだトップレベルで直接公開していない新しいリクエストフィールドが必要な場合は、`extra_args` を使用します。
|
||||
|
||||
また、OpenAI の Responses API を使用する場合、[他にもいくつかの任意パラメーター](https://platform.openai.com/docs/api-reference/responses/create)(例: `user`、`service_tier` など)があります。それらがトップレベルで利用できない場合は、`extra_args` を使用して渡すこともできます。同じリクエストフィールドを、直接の `ModelSettings` フィールドでも同時に設定しないでください。
|
||||
|
||||
```python
|
||||
from agents import Agent, ModelSettings
|
||||
|
||||
english_agent = Agent(
|
||||
name="English agent",
|
||||
instructions="You only speak English",
|
||||
model="gpt-4.1",
|
||||
model_settings=ModelSettings(
|
||||
temperature=0.1,
|
||||
extra_args={"service_tier": "flex", "user": "user_12345"},
|
||||
),
|
||||
)
|
||||
```
|
||||
|
||||
## Runner 管理の再試行
|
||||
|
||||
再試行はランタイム専用で、明示的な有効化が必要です。`ModelSettings(retry=...)` を設定し、再試行ポリシーが再試行を選択しない限り、SDK は一般的なモデルリクエストを再試行しません。
|
||||
|
||||
```python
|
||||
from agents import Agent, ModelRetrySettings, ModelSettings, retry_policies
|
||||
|
||||
agent = Agent(
|
||||
name="Assistant",
|
||||
model="gpt-5.5",
|
||||
model_settings=ModelSettings(
|
||||
retry=ModelRetrySettings(
|
||||
max_retries=4,
|
||||
backoff={
|
||||
"initial_delay": 0.5,
|
||||
"max_delay": 5.0,
|
||||
"multiplier": 2.0,
|
||||
"jitter": True,
|
||||
},
|
||||
policy=retry_policies.any(
|
||||
retry_policies.provider_suggested(),
|
||||
retry_policies.retry_after(),
|
||||
retry_policies.network_error(),
|
||||
retry_policies.http_status([408, 409, 429, 500, 502, 503, 504]),
|
||||
),
|
||||
)
|
||||
),
|
||||
)
|
||||
```
|
||||
|
||||
`ModelRetrySettings` には 3 つのフィールドがあります。
|
||||
|
||||
<div class="field-table" markdown="1">
|
||||
|
||||
| フィールド | 型 | 注記 |
|
||||
| --- | --- | --- |
|
||||
| `max_retries` | `int | None` | 初回リクエスト後に許可される再試行回数。 |
|
||||
| `backoff` | `ModelRetryBackoffSettings | dict | None` | ポリシーが明示的な遅延を返さずに再試行する場合のデフォルトの遅延戦略。`backoff.max_delay` は、この計算された backoff 遅延のみを上限設定します。ポリシーから返された明示的な遅延や retry-after ヒントには上限を設定しません。 |
|
||||
| `policy` | `RetryPolicy | None` | 再試行するかどうかを決定するコールバック。このフィールドはランタイム専用で、シリアライズされません。 |
|
||||
|
||||
</div>
|
||||
|
||||
再試行ポリシーは、以下を含む [`RetryPolicyContext`][agents.retry.RetryPolicyContext] を受け取ります。
|
||||
|
||||
- `attempt` と `max_retries`。試行回数を考慮した判断を行えます。
|
||||
- `stream`。ストリーミングと非ストリーミングの動作を分岐できます。
|
||||
- `error`。生の内容を確認できます。
|
||||
- `status_code`、`retry_after`、`error_code`、`is_network_error`、`is_timeout`、`is_abort` などの正規化された情報。
|
||||
- 基盤のモデルアダプターが再試行ガイダンスを提供できる場合の `provider_advice`。
|
||||
|
||||
ポリシーは以下のいずれかを返せます。
|
||||
|
||||
- 単純な再試行判断のための `True` / `False`。
|
||||
- 遅延を上書きしたり、診断理由を付加したりしたい場合の [`RetryDecision`][agents.retry.RetryDecision]。
|
||||
|
||||
SDK は、`retry_policies` に既製のヘルパーをエクスポートしています。
|
||||
|
||||
| ヘルパー | 動作 |
|
||||
| --- | --- |
|
||||
| `retry_policies.never()` | 常にオプトアウトします。 |
|
||||
| `retry_policies.provider_suggested()` | 利用可能な場合、プロバイダーの再試行助言に従います。 |
|
||||
| `retry_policies.network_error()` | 一時的なトランスポート障害とタイムアウト障害に一致します。 |
|
||||
| `retry_policies.http_status([...])` | 選択した HTTP ステータスコードに一致します。 |
|
||||
| `retry_policies.retry_after()` | retry-after ヒントが利用可能な場合にのみ、その遅延を使用して再試行します。このヘルパーは retry-after 値を明示的なポリシー遅延として扱うため、`backoff.max_delay` はそれを上限設定しません。 |
|
||||
| `retry_policies.any(...)` | 入れ子のポリシーのいずれかが有効化した場合に再試行します。 |
|
||||
| `retry_policies.all(...)` | 入れ子のすべてのポリシーが有効化した場合にのみ再試行します。 |
|
||||
|
||||
ポリシーを合成する場合、`provider_suggested()` は最も安全な最初の構成要素です。プロバイダーがそれらを区別できる場合に、プロバイダーの拒否とリプレイ安全性の承認を保持するためです。
|
||||
|
||||
##### 安全境界
|
||||
|
||||
一部の障害は自動的には再試行されません。
|
||||
|
||||
- 中止エラー。
|
||||
- プロバイダーからの助言がリプレイを安全でないと示すリクエスト。
|
||||
- リプレイが安全でなくなる形で出力がすでに開始された後のストリーミング実行。
|
||||
|
||||
`previous_response_id` または `conversation_id` を使用するステートフルな後続リクエストも、より保守的に扱われます。これらのリクエストでは、`network_error()` や `http_status([500])` など、プロバイダー以外の述語だけでは十分ではありません。再試行ポリシーには、通常 `retry_policies.provider_suggested()` を通じて、プロバイダーからのリプレイ安全性の承認を含める必要があります。
|
||||
|
||||
##### Runner とエージェントのマージ動作
|
||||
|
||||
`retry` は、runner レベルとエージェントレベルの `ModelSettings` の間でディープマージされます。
|
||||
|
||||
- エージェントは `retry.max_retries` だけを上書きし、runner の `policy` を継承できます。
|
||||
- エージェントは `retry.backoff` の一部だけを上書きし、runner の兄弟 backoff フィールドを維持できます。
|
||||
- `policy` はランタイム専用であるため、シリアライズされた `ModelSettings` は `max_retries` と `backoff` を保持しますが、コールバック自体は省略します。
|
||||
|
||||
より完全な例については、[`examples/basic/retry.py`](https://github.com/openai/openai-agents-python/tree/main/examples/basic/retry.py) と [アダプターベースの再試行例](https://github.com/openai/openai-agents-python/tree/main/examples/basic/retry_litellm.py)を参照してください。
|
||||
|
||||
## 非 OpenAI プロバイダーのトラブルシューティング
|
||||
|
||||
### トレーシングクライアントエラー 401
|
||||
|
||||
トレーシングに関連するエラーが発生する場合、トレースが OpenAI サーバーにアップロードされる一方で、OpenAI API キーがないことが原因です。これを解決するには、3 つの選択肢があります。
|
||||
|
||||
1. トレーシングを完全に無効にする: [`set_tracing_disabled(True)`][agents.set_tracing_disabled]。
|
||||
2. トレーシング用の OpenAI キーを設定する: [`set_tracing_export_api_key(...)`][agents.set_tracing_export_api_key]。この API キーはトレースのアップロードにのみ使用され、[platform.openai.com](https://platform.openai.com/) のものである必要があります。
|
||||
3. 非 OpenAI トレースプロセッサーを使用する。[トレーシングドキュメント](../tracing.md#custom-tracing-processors)を参照してください。
|
||||
|
||||
### Responses API のサポート
|
||||
|
||||
SDK はデフォルトで Responses API を使用しますが、他の多くの LLM プロバイダーはまだこれをサポートしていません。その結果、404 などの問題が発生する場合があります。解決するには、2 つの選択肢があります。
|
||||
|
||||
1. [`set_default_openai_api("chat_completions")`][agents.set_default_openai_api] を呼び出す。これは、環境変数で `OPENAI_API_KEY` と `OPENAI_BASE_URL` を設定している場合に機能します。
|
||||
2. [`OpenAIChatCompletionsModel`][agents.models.openai_chatcompletions.OpenAIChatCompletionsModel] を使用する。[こちら](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/)に例があります。
|
||||
|
||||
### structured outputs のサポート
|
||||
|
||||
一部のモデルプロバイダーは、[structured outputs](https://platform.openai.com/docs/guides/structured-outputs) をサポートしていません。これにより、次のようなエラーが発生する場合があります。
|
||||
|
||||
```
|
||||
|
||||
BadRequestError: Error code: 400 - {'error': {'message': "'response_format.type' : value is not one of the allowed values ['text','json_object']", 'type': 'invalid_request_error'}}
|
||||
|
||||
```
|
||||
|
||||
これは一部のモデルプロバイダーの制約です。JSON 出力はサポートしていますが、出力に使用する `json_schema` の指定は許可していません。この問題の修正に取り組んでいますが、JSON schema 出力をサポートしているプロバイダーに依存することを推奨します。そうしないと、不正な形式の JSON が原因でアプリが頻繁に壊れるためです。
|
||||
|
||||
## プロバイダー間でのモデルの混在
|
||||
|
||||
モデルプロバイダー間の機能差を理解しておく必要があります。そうしないと、エラーが発生する可能性があります。たとえば、OpenAI は structured outputs、マルチモーダル入力、ホスト型のファイル検索と Web 検索をサポートしていますが、他の多くのプロバイダーはこれらの機能をサポートしていません。次の制限に注意してください。
|
||||
|
||||
- サポートされていない `tools` を、それらを理解しないプロバイダーに送信しないでください
|
||||
- テキスト専用のモデルを呼び出す前に、マルチモーダル入力を除外してください
|
||||
- structured JSON 出力をサポートしていないプロバイダーは、無効な JSON を生成することがある点に注意してください。
|
||||
|
||||
## サードパーティアダプター
|
||||
|
||||
SDK の組み込みプロバイダー統合ポイントで十分でない場合にのみ、サードパーティアダプターを使用してください。この SDK で OpenAI モデルのみを使用している場合は、Any-LLM や LiteLLM ではなく、組み込みの [`OpenAIResponsesModel`][agents.models.openai_responses.OpenAIResponsesModel] パスを優先してください。サードパーティアダプターは、OpenAI モデルと非 OpenAI プロバイダーを組み合わせる必要がある場合、または組み込みパスでは提供されないアダプター管理のプロバイダー対応範囲やルーティングが必要な場合のためのものです。アダプターは SDK と上流のモデルプロバイダーの間に別の互換性レイヤーを追加するため、機能サポートとリクエストセマンティクスはプロバイダーによって異なる場合があります。SDK には現在、Any-LLM と LiteLLM がベストエフォートのベータ版アダプター統合として含まれています。
|
||||
|
||||
### Any-LLM
|
||||
|
||||
Any-LLM のサポートは、Any-LLM 管理のプロバイダー対応範囲やルーティングが必要な場合向けに、ベストエフォートのベータ版として含まれています。
|
||||
|
||||
上流プロバイダーパスによって、Any-LLM は Responses API、Chat Completions 互換 API、またはプロバイダー固有の互換性レイヤーを使用する場合があります。
|
||||
|
||||
Any-LLM が必要な場合は、`openai-agents[any-llm]` をインストールし、[`examples/model_providers/any_llm_auto.py`](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/any_llm_auto.py) または [`examples/model_providers/any_llm_provider.py`](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/any_llm_provider.py) から始めてください。[`MultiProvider`][agents.MultiProvider] で `any-llm/...` モデル名を使用するか、`AnyLLMModel` を直接インスタンス化するか、実行スコープで `AnyLLMProvider` を使用できます。モデルサーフェスを明示的に固定する必要がある場合は、`AnyLLMModel` を構築するときに `api="responses"` または `api="chat_completions"` を渡します。
|
||||
|
||||
Any-LLM はサードパーティアダプターレイヤーであり続けるため、プロバイダー依存関係と機能ギャップは SDK ではなく Any-LLM によって上流で定義されます。使用状況メトリクスは、上流プロバイダーが返す場合に自動的に伝播されますが、ストリーミング Chat Completions バックエンドでは、usage chunks を出力する前に `ModelSettings(include_usage=True)` が必要になる場合があります。structured outputs、ツール呼び出し、使用状況レポート、または Responses 固有の動作に依存する場合は、デプロイ予定の正確なプロバイダーバックエンドを検証してください。
|
||||
|
||||
### LiteLLM
|
||||
|
||||
LiteLLM のサポートは、LiteLLM 固有のプロバイダー対応範囲やルーティングが必要な場合向けに、ベストエフォートのベータ版として含まれています。
|
||||
|
||||
LiteLLM が必要な場合は、`openai-agents[litellm]` をインストールし、[`examples/model_providers/litellm_auto.py`](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/litellm_auto.py) または [`examples/model_providers/litellm_provider.py`](https://github.com/openai/openai-agents-python/tree/main/examples/model_providers/litellm_provider.py) から始めてください。`litellm/...` モデル名を使用するか、[`LitellmModel`][agents.extensions.models.litellm_model.LitellmModel] を直接インスタンス化できます。
|
||||
|
||||
一部の LiteLLM ベースのプロバイダーは、デフォルトでは SDK の使用状況メトリクスを設定しません。使用状況レポートが必要な場合は、`ModelSettings(include_usage=True)` を渡し、structured outputs、ツール呼び出し、使用状況レポート、またはアダプター固有のルーティング動作に依存する場合は、デプロイ予定の正確なプロバイダーバックエンドを検証してください。
|
||||
@@ -0,0 +1,13 @@
|
||||
---
|
||||
search:
|
||||
exclude: true
|
||||
---
|
||||
# LiteLLM
|
||||
|
||||
<script>
|
||||
window.location.replace("../#third-party-adapters");
|
||||
</script>
|
||||
|
||||
このページは [Models のサードパーティアダプターセクション](index.md#third-party-adapters) に移動しました。
|
||||
|
||||
自動的にリダイレクトされない場合は、上記のリンクを使用してください。
|
||||
+50
-23
@@ -1,37 +1,64 @@
|
||||
# 複数のエージェントのオーケストレーション
|
||||
---
|
||||
search:
|
||||
exclude: true
|
||||
---
|
||||
# エージェントオーケストレーション
|
||||
|
||||
オーケストレーションとは、アプリ内でのエージェントの流れを指します。どのエージェントが、どの順番で実行され、次に何が起こるかをどのように決定するかということです。エージェントをオーケストレーションする主な方法は 2 つあります。
|
||||
オーケストレーションとは、アプリ内でのエージェントの流れを指します。どのエージェントを、どの順序で実行し、次に何が起こるかをどのように決定するか、ということです。エージェントをオーケストレーションする主な方法は 2 つあります:
|
||||
|
||||
1. LLM に意思決定を任せる方法:これは LLM の知能を活用し、計画・推論・意思決定を行わせるものです。
|
||||
2. コードによるオーケストレーション:コードでエージェントの流れを制御する方法です。
|
||||
1. LLM に判断を任せる:LLM の知能を使って計画と推論を行い、それに基づいて取る手順を決定します。
|
||||
2. コードによるオーケストレーション:コードでエージェントの流れを決定します。
|
||||
|
||||
これらのパターンは組み合わせて使うこともできます。それぞれにメリット・デメリットがあり、以下で説明します。
|
||||
これらのパターンは組み合わせることができます。それぞれにトレードオフがあり、以下で説明します。
|
||||
|
||||
## LLM によるオーケストレーション
|
||||
|
||||
エージェントとは、instructions、tools、handoffs を備えた LLM です。つまり、オープンエンドなタスクが与えられた場合、LLM は自律的にタスクへの取り組み方を計画し、tools を使ってアクションを実行・データを取得し、handoffs を使ってサブエージェントにタスクを委任できます。たとえば、リサーチエージェントには以下のような tools を持たせることができます。
|
||||
エージェントとは、instructions、tools、ハンドオフを備えた LLM です。つまり、自由度の高いタスクが与えられた場合、LLM は、tools を使ってアクションを実行しデータを取得し、ハンドオフを使ってサブエージェントにタスクを委任しながら、そのタスクへの取り組み方を自律的に計画できます。たとえば、リサーチエージェントには次のようなツールを備えられます:
|
||||
|
||||
- Web 検索でオンライン情報を探す
|
||||
- ファイル検索やリトリーバルで独自データや接続先を検索する
|
||||
- コンピュータ操作でコンピュータ上のアクションを実行する
|
||||
- コード実行でデータ分析を行う
|
||||
- 計画やレポート作成などに特化したエージェントへの handoffs
|
||||
- オンラインで情報を見つけるための Web 検索
|
||||
- 独自データや接続先を検索するためのファイル検索と取得
|
||||
- コンピューター上でアクションを実行するためのコンピュータ操作
|
||||
- データ分析を行うためのコード実行
|
||||
- 計画、レポート作成などに優れた専門エージェントへのハンドオフ。
|
||||
|
||||
このパターンは、タスクがオープンエンドで LLM の知能に頼りたい場合に最適です。ここで重要なポイントは次のとおりです。
|
||||
### コア SDK パターン
|
||||
|
||||
1. 良いプロンプトに投資しましょう。利用可能な tools、使い方、守るべきパラメーターを明確に伝えます。
|
||||
2. アプリをモニタリングし、改善を繰り返しましょう。問題が発生した箇所を確認し、プロンプトを改善します。
|
||||
3. エージェントに内省と改善を促しましょう。たとえばループで実行し、自己批評させたり、エラーメッセージを与えて改善させたりします。
|
||||
4. 何でもできる汎用エージェントではなく、1 つのタスクに特化したエージェントを用意しましょう。
|
||||
5. [evals](https://platform.openai.com/docs/guides/evals) に投資しましょう。これによりエージェントを訓練し、タスクの精度を向上させることができます。
|
||||
Python SDK では、次の 2 つのオーケストレーションパターンが最もよく登場します:
|
||||
|
||||
| パターン | 仕組み | 最適な場合 |
|
||||
| --- | --- | --- |
|
||||
| Agents as tools | マネージャーエージェントが会話の制御を維持し、`Agent.as_tool()` を通じて専門エージェントを呼び出します。 | 1 つのエージェントに最終回答を担わせたい場合、複数の専門エージェントからの出力を統合したい場合、または共通のガードレールを 1 か所で適用したい場合。 |
|
||||
| ハンドオフ | トリアージエージェントが会話を専門エージェントにルーティングし、その専門エージェントがそのターンの残りでアクティブなエージェントになります。 | 専門エージェントに直接応答させたい場合、プロンプトを焦点の絞られた状態に保ちたい場合、またはマネージャーが結果を説明することなく instructions を切り替えたい場合。 |
|
||||
|
||||
専門エージェントが範囲の限定されたサブタスクを支援するべきだが、ユーザー向けの会話を引き継ぐべきではない場合は、 **agents as tools** を使用します。ルーティング自体がワークフローの一部であり、選ばれた専門エージェントにインタラクションの次の部分を担わせたい場合は、 **ハンドオフ** を使用します。
|
||||
|
||||
この 2 つを組み合わせることもできます。トリアージエージェントが専門エージェントにハンドオフし、その専門エージェントがさらに狭いサブタスクのために他のエージェントをツールとして呼び出すこともできます。
|
||||
|
||||
このパターンは、タスクの自由度が高く、LLM の知能に頼りたい場合に最適です。ここで最も重要な戦術は次のとおりです:
|
||||
|
||||
1. 優れたプロンプトに投資します。利用できるツール、その使い方、そしてエージェントが従うべきパラメーターを明確にします。
|
||||
2. アプリを監視し、反復改善します。どこで問題が起こるかを確認し、プロンプトを改善します。
|
||||
3. エージェントが内省して改善できるようにします。たとえば、ループ内で実行して自己批評させる、またはエラーメッセージを提供して改善させます。
|
||||
4. 何でも得意であることを期待される汎用エージェントではなく、1 つのタスクに秀でた専門エージェントを用意します。
|
||||
5. [evals](https://platform.openai.com/docs/guides/evals) に投資します。これにより、エージェントをトレーニングして改善し、タスクの遂行能力を高めることができます。
|
||||
|
||||
このスタイルのオーケストレーションを支えるコア SDK の基本コンポーネントを知りたい場合は、[ツール](tools.md)、[ハンドオフ](handoffs.md)、[エージェントの実行](running_agents.md) から始めてください。
|
||||
|
||||
## コードによるオーケストレーション
|
||||
|
||||
LLM によるオーケストレーションは強力ですが、コードによるオーケストレーションは、速度・コスト・パフォーマンスの面でより決定的かつ予測可能になります。よく使われるパターンは次のとおりです。
|
||||
LLM によるオーケストレーションは強力ですが、コードによるオーケストレーションは、速度、コスト、パフォーマンスの観点でタスクをより決定論的で予測可能にします。ここでの一般的なパターンは次のとおりです:
|
||||
|
||||
- [structured outputs](https://platform.openai.com/docs/guides/structured-outputs) を使い、コードで検査できる適切な形式のデータを生成する。たとえば、エージェントにタスクをいくつかのカテゴリーに分類させ、そのカテゴリーに応じて次のエージェントを選択することができます。
|
||||
- 複数のエージェントを連鎖させ、一方の出力を次の入力に変換する。たとえば、ブログ記事の作成タスクを「リサーチ→アウトライン作成→記事執筆→批評→改善」といった一連のステップに分解できます。
|
||||
- タスクを実行するエージェントと、評価・フィードバックを行うエージェントを `while` ループで回し、評価者が基準を満たしたと判断するまで繰り返します。
|
||||
- 複数のエージェントを並列で実行する(例:Python の基本コンポーネントである `asyncio.gather` を利用)。これは、互いに依存しない複数のタスクを高速に処理したい場合に有効です。
|
||||
- [structured outputs](https://platform.openai.com/docs/guides/structured-outputs) を使って、コードで検査できる適切な形式のデータを生成します。たとえば、エージェントにタスクをいくつかのカテゴリーに分類させ、そのカテゴリーに基づいて次のエージェントを選択できます。
|
||||
- 複数のエージェントをチェーンし、あるエージェントの出力を次のエージェントの入力に変換します。ブログ記事を書くようなタスクを一連のステップに分解できます - リサーチする、アウトラインを書く、ブログ記事を書く、批評し、それから改善します。
|
||||
- タスクを実行するエージェントを `while` ループ内で、評価してフィードバックを提供するエージェントと一緒に実行し、評価者が出力が特定の基準を満たしたと言うまで続けます。
|
||||
- 複数のエージェントを並列に実行します。たとえば、`asyncio.gather` のような Python の基本コンポーネントを使います。これは、互いに依存しない複数のタスクがある場合に高速化に役立ちます。
|
||||
|
||||
[`examples/agent_patterns`](https://github.com/openai/openai-agents-python/tree/main/examples/agent_patterns) には、さまざまな code examples をご用意しています。
|
||||
[`examples/agent_patterns`](https://github.com/openai/openai-agents-python/tree/main/examples/agent_patterns) には多数のコード例があります。
|
||||
|
||||
## 関連ガイド
|
||||
|
||||
- 構成パターンとエージェント設定については、[エージェント](agents.md) を参照してください。
|
||||
- `Agent.as_tool()` とマネージャースタイルのオーケストレーションについては、[ツール](tools.md#agents-as-tools) を参照してください。
|
||||
- 専門エージェント間の委任については、[ハンドオフ](handoffs.md) を参照してください。
|
||||
- 実行ごとのオーケストレーション制御と会話状態については、[エージェントの実行](running_agents.md) を参照してください。
|
||||
- 最小限のエンドツーエンドのハンドオフ例については、[クイックスタート](quickstart.md) を参照してください。
|
||||
+164
-128
@@ -1,8 +1,12 @@
|
||||
---
|
||||
search:
|
||||
exclude: true
|
||||
---
|
||||
# クイックスタート
|
||||
|
||||
## プロジェクトと仮想環境の作成
|
||||
|
||||
この作業は一度だけ行えば十分です。
|
||||
これは一度だけ行えば十分です。
|
||||
|
||||
```bash
|
||||
mkdir my_project
|
||||
@@ -12,12 +16,20 @@ python -m venv .venv
|
||||
|
||||
### 仮想環境の有効化
|
||||
|
||||
新しいターミナルセッションを開始するたびに、これを実行してください。
|
||||
新しいターミナルセッションを開始するたびに行ってください。
|
||||
|
||||
macOS または Linux の場合:
|
||||
|
||||
```bash
|
||||
source .venv/bin/activate
|
||||
```
|
||||
|
||||
Windows の場合:
|
||||
|
||||
```cmd
|
||||
.venv\Scripts\activate
|
||||
```
|
||||
|
||||
### Agents SDK のインストール
|
||||
|
||||
```bash
|
||||
@@ -26,164 +38,188 @@ pip install openai-agents # or `uv add openai-agents`, etc
|
||||
|
||||
### OpenAI API キーの設定
|
||||
|
||||
まだお持ちでない場合は、[こちらの手順](https://platform.openai.com/docs/quickstart#create-and-export-an-api-key)に従って OpenAI API キーを作成してください。
|
||||
まだ持っていない場合は、[こちらの手順](https://platform.openai.com/docs/quickstart#create-and-export-an-api-key)に従って OpenAI API キーを作成してください。
|
||||
|
||||
以下のコマンドは、現在のターミナルセッションにキーを設定します。
|
||||
|
||||
macOS または Linux の場合:
|
||||
|
||||
```bash
|
||||
export OPENAI_API_KEY=sk-...
|
||||
```
|
||||
|
||||
## 最初のエージェントを作成する
|
||||
Windows PowerShell の場合:
|
||||
|
||||
エージェントは instructions、名前、およびオプションの設定(例: `model_config` )で定義します。
|
||||
```powershell
|
||||
$env:OPENAI_API_KEY = "sk-..."
|
||||
```
|
||||
|
||||
Windows コマンドプロンプトの場合:
|
||||
|
||||
```cmd
|
||||
set "OPENAI_API_KEY=sk-..."
|
||||
```
|
||||
|
||||
## 最初のエージェントの作成
|
||||
|
||||
エージェントは、instructions、名前、および特定のモデルなどの任意の設定で定義します。
|
||||
|
||||
```python
|
||||
from agents import Agent
|
||||
|
||||
agent = Agent(
|
||||
name="Math Tutor",
|
||||
instructions="You provide help with math problems. Explain your reasoning at each step and include examples",
|
||||
)
|
||||
```
|
||||
|
||||
## さらにエージェントを追加する
|
||||
|
||||
追加のエージェントも同様の方法で定義できます。 `handoff_descriptions` は、ハンドオフのルーティングを判断するための追加コンテキストを提供します。
|
||||
|
||||
```python
|
||||
from agents import Agent
|
||||
|
||||
history_tutor_agent = Agent(
|
||||
name="History Tutor",
|
||||
handoff_description="Specialist agent for historical questions",
|
||||
instructions="You provide assistance with historical queries. Explain important events and context clearly.",
|
||||
)
|
||||
|
||||
math_tutor_agent = Agent(
|
||||
name="Math Tutor",
|
||||
handoff_description="Specialist agent for math questions",
|
||||
instructions="You provide help with math problems. Explain your reasoning at each step and include examples",
|
||||
instructions="You answer history questions clearly and concisely.",
|
||||
)
|
||||
```
|
||||
|
||||
## ハンドオフを定義する
|
||||
## 最初のエージェントの実行
|
||||
|
||||
各エージェントで、タスクを進めるために選択できる送信先ハンドオフオプションのインベントリを定義できます。
|
||||
[`Runner`][agents.run.Runner] を使用してエージェントを実行し、[`RunResult`][agents.result.RunResult] を取得します。
|
||||
|
||||
```python
|
||||
triage_agent = Agent(
|
||||
name="Triage Agent",
|
||||
instructions="You determine which agent to use based on the user's homework question",
|
||||
handoffs=[history_tutor_agent, math_tutor_agent]
|
||||
)
|
||||
```
|
||||
|
||||
## エージェントのオーケストレーションを実行する
|
||||
|
||||
ワークフローが正しく動作し、トリアージエージェントが 2 つの専門エージェント間で正しくルーティングするか確認しましょう。
|
||||
|
||||
```python
|
||||
from agents import Runner
|
||||
|
||||
async def main():
|
||||
result = await Runner.run(triage_agent, "What is the capital of France?")
|
||||
print(result.final_output)
|
||||
```
|
||||
|
||||
## ガードレールを追加する
|
||||
|
||||
入力または出力に対して実行するカスタムガードレールを定義できます。
|
||||
|
||||
```python
|
||||
from agents import GuardrailFunctionOutput, Agent, Runner
|
||||
from pydantic import BaseModel
|
||||
|
||||
class HomeworkOutput(BaseModel):
|
||||
is_homework: bool
|
||||
reasoning: str
|
||||
|
||||
guardrail_agent = Agent(
|
||||
name="Guardrail check",
|
||||
instructions="Check if the user is asking about homework.",
|
||||
output_type=HomeworkOutput,
|
||||
)
|
||||
|
||||
async def homework_guardrail(ctx, agent, input_data):
|
||||
result = await Runner.run(guardrail_agent, input_data, context=ctx.context)
|
||||
final_output = result.final_output_as(HomeworkOutput)
|
||||
return GuardrailFunctionOutput(
|
||||
output_info=final_output,
|
||||
tripwire_triggered=not final_output.is_homework,
|
||||
)
|
||||
```
|
||||
|
||||
## すべてをまとめる
|
||||
|
||||
すべてを組み合わせて、ハンドオフと入力ガードレールを使ったワークフロー全体を実行してみましょう。
|
||||
|
||||
```python
|
||||
from agents import Agent, InputGuardrail, GuardrailFunctionOutput, Runner
|
||||
from pydantic import BaseModel
|
||||
import asyncio
|
||||
from agents import Agent, Runner
|
||||
|
||||
class HomeworkOutput(BaseModel):
|
||||
is_homework: bool
|
||||
reasoning: str
|
||||
|
||||
guardrail_agent = Agent(
|
||||
name="Guardrail check",
|
||||
instructions="Check if the user is asking about homework.",
|
||||
output_type=HomeworkOutput,
|
||||
)
|
||||
|
||||
math_tutor_agent = Agent(
|
||||
name="Math Tutor",
|
||||
handoff_description="Specialist agent for math questions",
|
||||
instructions="You provide help with math problems. Explain your reasoning at each step and include examples",
|
||||
)
|
||||
|
||||
history_tutor_agent = Agent(
|
||||
agent = Agent(
|
||||
name="History Tutor",
|
||||
handoff_description="Specialist agent for historical questions",
|
||||
instructions="You provide assistance with historical queries. Explain important events and context clearly.",
|
||||
)
|
||||
|
||||
|
||||
async def homework_guardrail(ctx, agent, input_data):
|
||||
result = await Runner.run(guardrail_agent, input_data, context=ctx.context)
|
||||
final_output = result.final_output_as(HomeworkOutput)
|
||||
return GuardrailFunctionOutput(
|
||||
output_info=final_output,
|
||||
tripwire_triggered=not final_output.is_homework,
|
||||
)
|
||||
|
||||
triage_agent = Agent(
|
||||
name="Triage Agent",
|
||||
instructions="You determine which agent to use based on the user's homework question",
|
||||
handoffs=[history_tutor_agent, math_tutor_agent],
|
||||
input_guardrails=[
|
||||
InputGuardrail(guardrail_function=homework_guardrail),
|
||||
],
|
||||
instructions="You answer history questions clearly and concisely.",
|
||||
)
|
||||
|
||||
async def main():
|
||||
result = await Runner.run(triage_agent, "who was the first president of the united states?")
|
||||
print(result.final_output)
|
||||
|
||||
result = await Runner.run(triage_agent, "what is life")
|
||||
result = await Runner.run(agent, "When did the Roman Empire fall?")
|
||||
print(result.final_output)
|
||||
|
||||
if __name__ == "__main__":
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
## トレースを確認する
|
||||
2 回目のターンでは、`result.to_input_list()` を `Runner.run(...)` に戻して渡すか、[セッション](sessions/index.md)をアタッチするか、`conversation_id` / `previous_response_id` で OpenAI によりサーバー側で管理される状態を再利用できます。[エージェントの実行](running_agents.md)ガイドでは、これらのアプローチを比較しています。
|
||||
|
||||
エージェントの実行中に何が起こったかを確認するには、[OpenAI ダッシュボードの Trace viewer](https://platform.openai.com/traces) にアクセスして、エージェント実行のトレースを表示してください。
|
||||
次の目安を使ってください:
|
||||
|
||||
| 実現したいこと... | まず使うもの... |
|
||||
| --- | --- |
|
||||
| 完全な手動制御とプロバイダー非依存の履歴 | `result.to_input_list()` |
|
||||
| SDK に履歴の読み込みと保存を任せる | [`session=...`](sessions/index.md) |
|
||||
| OpenAI 管理のサーバー側継続 | `previous_response_id` または `conversation_id` |
|
||||
|
||||
トレードオフと正確な挙動については、[エージェントの実行](running_agents.md#choose-a-memory-strategy)を参照してください。
|
||||
|
||||
タスクが主にプロンプト、ツール、会話状態で完結する場合は、シンプルな `Agent` と `Runner` を使います。エージェントが分離されたワークスペース内の実ファイルを検査または変更する必要がある場合は、[Sandbox エージェントのクイックスタート](sandbox_agents.md)に進んでください。
|
||||
|
||||
## エージェントへのツールの付与
|
||||
|
||||
エージェントにツールを与えることで、情報を調べたりアクションを実行したりできます。
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
from agents import Agent, Runner, function_tool
|
||||
|
||||
|
||||
@function_tool
|
||||
def history_fun_fact() -> str:
|
||||
"""Return a short history fact."""
|
||||
return "Sharks are older than trees."
|
||||
|
||||
|
||||
agent = Agent(
|
||||
name="History Tutor",
|
||||
instructions="Answer history questions clearly. Use history_fun_fact when it helps.",
|
||||
tools=[history_fun_fact],
|
||||
)
|
||||
|
||||
|
||||
async def main():
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
"Tell me something surprising about ancient life on Earth.",
|
||||
)
|
||||
print(result.final_output)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
## さらにいくつかのエージェントの追加
|
||||
|
||||
マルチエージェントパターンを選ぶ前に、最終回答の主導権を誰が持つべきかを決めてください。
|
||||
|
||||
- **ハンドオフ**: スペシャリストが、そのターンの該当部分について会話を引き継ぎます。
|
||||
- **Agents as tools**: オーケストレーターが制御を維持し、スペシャリストをツールとして呼び出します。
|
||||
|
||||
このクイックスタートでは、最初の例として最も短いため、 **ハンドオフ** で続けます。マネージャースタイルのパターンについては、[エージェントオーケストレーション](multi_agent.md)と[ツール: agents as tools](tools.md#agents-as-tools)を参照してください。
|
||||
|
||||
追加のエージェントも同じ方法で定義できます。`handoff_description` は、いつ委譲すべきかについて、ルーティングエージェントに追加のコンテキストを提供します。
|
||||
|
||||
```python
|
||||
from agents import Agent
|
||||
|
||||
history_tutor_agent = Agent(
|
||||
name="History Tutor",
|
||||
handoff_description="Specialist agent for historical questions",
|
||||
instructions="You answer history questions clearly and concisely.",
|
||||
)
|
||||
|
||||
math_tutor_agent = Agent(
|
||||
name="Math Tutor",
|
||||
handoff_description="Specialist agent for math questions",
|
||||
instructions="You explain math step by step and include worked examples.",
|
||||
)
|
||||
```
|
||||
|
||||
## ハンドオフの定義
|
||||
|
||||
エージェントには、タスクを解決する際に選択できるハンドオフ先の選択肢の一覧を定義できます。
|
||||
|
||||
```python
|
||||
triage_agent = Agent(
|
||||
name="Triage Agent",
|
||||
instructions="Route each homework question to the right specialist.",
|
||||
handoffs=[history_tutor_agent, math_tutor_agent],
|
||||
)
|
||||
```
|
||||
|
||||
## エージェントオーケストレーションの実行
|
||||
|
||||
ランナーは、個々のエージェントの実行、すべてのハンドオフ、すべてのツール呼び出しを処理します。
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
from agents import Runner
|
||||
|
||||
|
||||
async def main():
|
||||
result = await Runner.run(
|
||||
triage_agent,
|
||||
"Who was the first president of the United States?",
|
||||
)
|
||||
print(result.final_output)
|
||||
print(f"Answered by: {result.last_agent.name}")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
## 参考コード例
|
||||
|
||||
このリポジトリには、同じ主要パターンに対応する完全なスクリプトが含まれています:
|
||||
|
||||
- [`examples/basic/hello_world.py`](https://github.com/openai/openai-agents-python/tree/main/examples/basic/hello_world.py) は最初の実行の例です。
|
||||
- [`examples/basic/tools.py`](https://github.com/openai/openai-agents-python/tree/main/examples/basic/tools.py) は関数ツールの例です。
|
||||
- [`examples/agent_patterns/routing.py`](https://github.com/openai/openai-agents-python/tree/main/examples/agent_patterns/routing.py) はマルチエージェントルーティングの例です。
|
||||
|
||||
## トレースの表示
|
||||
|
||||
エージェントの実行中に何が起きたかを確認するには、[OpenAI ダッシュボードのトレースビューアー](https://platform.openai.com/traces)に移動して、エージェント実行のトレースを表示してください。
|
||||
|
||||
## 次のステップ
|
||||
|
||||
より複雑なエージェントフローの構築方法を学びましょう:
|
||||
より複雑なエージェント型フローの構築方法を学びましょう:
|
||||
|
||||
- [エージェント](agents.md) の設定方法について学ぶ
|
||||
- [エージェントの実行](running_agents.md) について学ぶ
|
||||
- [tools](tools.md)、[guardrails](guardrails.md)、[models](models.md) について学ぶ
|
||||
- [エージェント](agents.md)の設定方法について学びます。
|
||||
- [エージェントの実行](running_agents.md)と[セッション](sessions/index.md)について学びます。
|
||||
- 作業を実際のワークスペース内で行う必要がある場合は、[Sandbox エージェント](sandbox_agents.md)について学びます。
|
||||
- [ツール](tools.md)、[ガードレール](guardrails.md)、[モデル](models/index.md)について学びます。
|
||||
@@ -0,0 +1,349 @@
|
||||
---
|
||||
search:
|
||||
exclude: true
|
||||
---
|
||||
# リアルタイムエージェントガイド
|
||||
|
||||
このガイドでは、OpenAI Agents SDK のリアルタイムレイヤーが OpenAI Realtime API にどのように対応するか、また Python SDK がその上にどのような追加動作を提供するかを説明します。
|
||||
|
||||
!!! warning "ベータ機能"
|
||||
|
||||
リアルタイムエージェントはベータ版です。実装を改善する中で、破壊的変更が入る可能性があります。
|
||||
|
||||
!!! note "はじめに"
|
||||
|
||||
デフォルトの Python パスを使いたい場合は、まず [クイックスタート](quickstart.md) を読んでください。アプリでサーバー側 WebSocket と SIP のどちらを使用すべきか判断している場合は、[Realtime トランスポート](transport.md) を読んでください。ブラウザーの WebRTC トランスポートは Python SDK の一部ではありません。
|
||||
|
||||
## 概要
|
||||
|
||||
リアルタイムエージェントは、モデルがテキストと音声を段階的に処理し、音声出力をストリーミングし、ツールを呼び出し、各ターンで新しいリクエストを再開始することなく中断を処理できるように、Realtime API への長時間接続を開いたままにします。
|
||||
|
||||
主な SDK コンポーネントは次のとおりです:
|
||||
|
||||
- **RealtimeAgent**: 1 つのリアルタイム専門エージェントに対する instructions、tools、出力ガードレール、ハンドオフ
|
||||
- **RealtimeRunner**: 開始エージェントをリアルタイムトランスポートに接続するセッションファクトリー
|
||||
- **RealtimeSession**: 入力を送信し、イベントを受信し、履歴を追跡し、ツールを実行するライブセッション
|
||||
- **RealtimeModel**: トランスポート抽象化。デフォルトは OpenAI のサーバー側 WebSocket 実装です。
|
||||
|
||||
## セッションライフサイクル
|
||||
|
||||
一般的なリアルタイムセッションは次のようになります:
|
||||
|
||||
1. `RealtimeAgent` を 1 つ以上作成します。
|
||||
2. 開始エージェントを指定して `RealtimeRunner` を作成します。
|
||||
3. `await runner.run()` を呼び出して `RealtimeSession` を取得します。
|
||||
4. `async with session:` または `await session.enter()` でセッションに入ります。
|
||||
5. `send_message()` または `send_audio()` でユーザー入力を送信します。
|
||||
6. 会話が終了するまでセッションイベントを反復処理します。
|
||||
|
||||
テキストのみの実行とは異なり、`runner.run()` はすぐに最終的な実行結果を生成しません。ローカル履歴、バックグラウンドのツール実行、ガードレール状態、アクティブなエージェント設定をトランスポート層と同期し続けるライブセッションオブジェクトを返します。
|
||||
|
||||
デフォルトでは、`RealtimeRunner` は `OpenAIRealtimeWebSocketModel` を使用するため、デフォルトの Python パスは Realtime API へのサーバー側 WebSocket 接続です。別の `RealtimeModel` を渡した場合でも、同じセッションライフサイクルとエージェント機能が適用され、接続の仕組みだけが変わることがあります。
|
||||
|
||||
## エージェントとセッションの設定
|
||||
|
||||
`RealtimeAgent` は、通常の `Agent` 型より意図的に範囲が絞られています:
|
||||
|
||||
- モデルの選択はエージェント単位ではなく、セッションレベルで設定します。
|
||||
- Structured outputs はサポートされていません。
|
||||
- 音声は設定できますが、セッションがすでに音声出力を生成した後は変更できません。
|
||||
- instructions、関数ツール、ハンドオフ、フック、出力ガードレールはすべて引き続き機能します。
|
||||
|
||||
`RealtimeSessionModelSettings` は、新しいネストされた `audio` 設定と、従来のフラットなエイリアスの両方をサポートします。新しいコードではネストされた形を推奨し、新しいリアルタイムエージェントでは `gpt-realtime-2` から始めてください:
|
||||
|
||||
```python
|
||||
runner = RealtimeRunner(
|
||||
starting_agent=agent,
|
||||
config={
|
||||
"model_settings": {
|
||||
"model_name": "gpt-realtime-2",
|
||||
"audio": {
|
||||
"input": {
|
||||
"format": "pcm16",
|
||||
"transcription": {"model": "gpt-4o-mini-transcribe"},
|
||||
"turn_detection": {"type": "semantic_vad", "interrupt_response": True},
|
||||
},
|
||||
"output": {"format": "pcm16", "voice": "ash"},
|
||||
},
|
||||
"tool_choice": "auto",
|
||||
}
|
||||
},
|
||||
)
|
||||
```
|
||||
|
||||
有用なセッションレベル設定には次のものがあります:
|
||||
|
||||
- `audio.input.format`, `audio.output.format`
|
||||
- `audio.input.transcription`
|
||||
- `audio.input.noise_reduction`
|
||||
- `audio.input.turn_detection`
|
||||
- `audio.output.voice`, `audio.output.speed`
|
||||
- `output_modalities`
|
||||
- `tool_choice`
|
||||
- `prompt`
|
||||
- `tracing`
|
||||
|
||||
`RealtimeRunner(config=...)` の有用な実行レベル設定には次のものがあります:
|
||||
|
||||
- `async_tool_calls`
|
||||
- `output_guardrails`
|
||||
- `guardrails_settings.debounce_text_length`
|
||||
- `tool_error_formatter`
|
||||
- `tracing_disabled`
|
||||
|
||||
型付きで利用できる全体のインターフェイスについては、[`RealtimeRunConfig`][agents.realtime.config.RealtimeRunConfig] と [`RealtimeSessionModelSettings`][agents.realtime.config.RealtimeSessionModelSettings] を参照してください。
|
||||
|
||||
## 入力と出力
|
||||
|
||||
### テキストと構造化されたユーザーメッセージ
|
||||
|
||||
プレーンテキストまたは構造化されたリアルタイムメッセージには、[`session.send_message()`][agents.realtime.session.RealtimeSession.send_message] を使用します。
|
||||
|
||||
```python
|
||||
from agents.realtime import RealtimeUserInputMessage
|
||||
|
||||
await session.send_message("Summarize what we discussed so far.")
|
||||
|
||||
message: RealtimeUserInputMessage = {
|
||||
"type": "message",
|
||||
"role": "user",
|
||||
"content": [
|
||||
{"type": "input_text", "text": "Describe this image."},
|
||||
{"type": "input_image", "image_url": image_data_url, "detail": "high"},
|
||||
],
|
||||
}
|
||||
await session.send_message(message)
|
||||
```
|
||||
|
||||
構造化メッセージは、リアルタイム会話に画像入力を含めるための主な方法です。[`examples/realtime/app/server.py`](https://github.com/openai/openai-agents-python/tree/main/examples/realtime/app/server.py) のサンプル Web デモでは、この方法で `input_image` メッセージを転送しています。
|
||||
|
||||
### 音声入力
|
||||
|
||||
未加工の音声バイトをストリーミングするには、[`session.send_audio()`][agents.realtime.session.RealtimeSession.send_audio] を使用します:
|
||||
|
||||
```python
|
||||
await session.send_audio(audio_bytes)
|
||||
```
|
||||
|
||||
サーバー側のターン検出が無効な場合、ターン境界をマークする責任はあなたにあります。高レベルの便利な方法は次のとおりです:
|
||||
|
||||
```python
|
||||
await session.send_audio(audio_bytes, commit=True)
|
||||
```
|
||||
|
||||
より低レベルの制御が必要な場合は、基盤となるモデルトランスポートを通じて `input_audio_buffer.commit` などの raw クライアントイベントを送信することもできます。
|
||||
|
||||
### 手動レスポンス制御
|
||||
|
||||
`session.send_message()` は高レベルのパスを使ってユーザー入力を送信し、レスポンスを開始します。raw 音声バッファリングは、すべての設定で同じことを自動的に行う **わけではありません**。
|
||||
|
||||
Realtime API レベルでは、手動ターン制御とは、raw の `session.update` で `turn_detection` をクリアし、その後 `input_audio_buffer.commit` と `response.create` を自分で送信することを意味します。
|
||||
|
||||
ターンを手動で管理している場合は、モデルトランスポートを通じて raw クライアントイベントを送信できます:
|
||||
|
||||
```python
|
||||
from agents.realtime.model_inputs import RealtimeModelSendRawMessage
|
||||
|
||||
await session.model.send_event(
|
||||
RealtimeModelSendRawMessage(
|
||||
message={
|
||||
"type": "response.create",
|
||||
}
|
||||
)
|
||||
)
|
||||
```
|
||||
|
||||
このパターンは次の場合に有用です:
|
||||
|
||||
- `turn_detection` が無効で、モデルがいつ応答すべきかを自分で決めたい場合
|
||||
- レスポンスをトリガーする前に、ユーザー入力を検査したり制御したりしたい場合
|
||||
- 帯域外レスポンス用のカスタムプロンプトが必要な場合
|
||||
|
||||
[`examples/realtime/twilio_sip/server.py`](https://github.com/openai/openai-agents-python/tree/main/examples/realtime/twilio_sip/server.py) の SIP 例では、冒頭の挨拶を強制するために raw の `response.create` を使用しています。
|
||||
|
||||
## イベント、履歴、中断
|
||||
|
||||
`RealtimeSession` は、必要に応じて raw モデルイベントも転送しつつ、より高レベルの SDK イベントを発行します。
|
||||
|
||||
特に有用なセッションイベントには次のものがあります:
|
||||
|
||||
- `audio`, `audio_end`, `audio_interrupted`
|
||||
- `agent_start`, `agent_end`
|
||||
- `tool_start`, `tool_end`, `tool_approval_required`
|
||||
- `handoff`
|
||||
- `history_added`, `history_updated`
|
||||
- `guardrail_tripped`
|
||||
- `input_audio_timeout_triggered`
|
||||
- `error`
|
||||
- `raw_model_event`
|
||||
|
||||
UI 状態に最も有用なイベントは通常 `history_added` と `history_updated` です。これらは、ユーザーメッセージ、アシスタントメッセージ、ツール呼び出しを含む、セッションのローカル履歴を `RealtimeItem` オブジェクトとして公開します。
|
||||
|
||||
### 中断と再生追跡
|
||||
|
||||
ユーザーがアシスタントを中断すると、セッションは `audio_interrupted` を発行し、サーバー側の会話がユーザーが実際に聞いた内容と一致するように履歴を更新します。
|
||||
|
||||
低遅延のローカル再生では、デフォルトの再生トラッカーで十分なことが多いです。リモート再生や遅延のある再生シナリオ、特にテレフォニーでは、生成された音声がすべてすでに聞かれたと仮定するのではなく、実際の再生進捗に基づいて中断時の切り詰めが行われるように [`RealtimePlaybackTracker`][agents.realtime.model.RealtimePlaybackTracker] を使用してください。
|
||||
|
||||
[`examples/realtime/twilio/twilio_handler.py`](https://github.com/openai/openai-agents-python/tree/main/examples/realtime/twilio/twilio_handler.py) の Twilio 例は、このパターンを示しています。
|
||||
|
||||
## ツール、承認、ハンドオフ、ガードレール
|
||||
|
||||
### 関数ツール
|
||||
|
||||
リアルタイムエージェントは、ライブ会話中に関数ツールをサポートします:
|
||||
|
||||
```python
|
||||
from agents import function_tool
|
||||
|
||||
|
||||
@function_tool
|
||||
def get_weather(city: str) -> str:
|
||||
"""Get current weather for a city."""
|
||||
return f"The weather in {city} is sunny, 72F."
|
||||
|
||||
|
||||
agent = RealtimeAgent(
|
||||
name="Assistant",
|
||||
instructions="You can answer weather questions.",
|
||||
tools=[get_weather],
|
||||
)
|
||||
```
|
||||
|
||||
### ツール承認
|
||||
|
||||
関数ツールは、実行前に人間による承認を必要とする場合があります。その場合、セッションは `tool_approval_required` を発行し、`approve_tool_call()` または `reject_tool_call()` を呼び出すまでツール実行を一時停止します。
|
||||
|
||||
```python
|
||||
async for event in session:
|
||||
if event.type == "tool_approval_required":
|
||||
await session.approve_tool_call(event.call_id)
|
||||
```
|
||||
|
||||
具体的なサーバー側の承認ループについては、[`examples/realtime/app/server.py`](https://github.com/openai/openai-agents-python/tree/main/examples/realtime/app/server.py) を参照してください。ヒューマンインザループのドキュメントでも、[ヒューマンインザループ](../human_in_the_loop.md) でこのフローを参照しています。
|
||||
|
||||
### ハンドオフ
|
||||
|
||||
リアルタイムハンドオフにより、あるエージェントがライブ会話を別の専門エージェントへ転送できます:
|
||||
|
||||
```python
|
||||
from agents.realtime import RealtimeAgent, realtime_handoff
|
||||
|
||||
billing_agent = RealtimeAgent(
|
||||
name="Billing Support",
|
||||
instructions="You specialize in billing issues.",
|
||||
)
|
||||
|
||||
main_agent = RealtimeAgent(
|
||||
name="Customer Service",
|
||||
instructions="Triage the request and hand off when needed.",
|
||||
handoffs=[realtime_handoff(billing_agent, tool_description="Transfer to billing support")],
|
||||
)
|
||||
```
|
||||
|
||||
`RealtimeAgent` をそのまま渡すハンドオフは自動的にラップされ、`realtime_handoff(...)` を使うと名前、説明、検証、コールバック、利用可否をカスタマイズできます。リアルタイムハンドオフは通常のハンドオフの `input_filter` をサポートしていません。
|
||||
|
||||
### ガードレール
|
||||
|
||||
リアルタイムエージェントでサポートされるのは出力ガードレールのみです。これらは部分トークンごとではなく、デバウンスされたトランスクリプトの蓄積に対して実行され、例外を発生させる代わりに `guardrail_tripped` を発行します。
|
||||
|
||||
```python
|
||||
from agents.guardrail import GuardrailFunctionOutput, OutputGuardrail
|
||||
|
||||
|
||||
def sensitive_data_check(context, agent, output):
|
||||
return GuardrailFunctionOutput(
|
||||
tripwire_triggered="password" in output,
|
||||
output_info=None,
|
||||
)
|
||||
|
||||
|
||||
agent = RealtimeAgent(
|
||||
name="Assistant",
|
||||
instructions="...",
|
||||
output_guardrails=[OutputGuardrail(guardrail_function=sensitive_data_check)],
|
||||
)
|
||||
```
|
||||
|
||||
リアルタイム出力ガードレールがトリップすると、セッションはアクティブなレスポンスを中断し、
|
||||
`response.cancel` を強制し、`guardrail_tripped` を発行し、トリガーされた
|
||||
ガードレールの名前を含むフォローアップのユーザーメッセージを送信するため、モデルは代替レスポンスを生成できます。音声プレイヤーは引き続き
|
||||
`audio_interrupted` をリッスンしてローカル再生を即座に停止する必要があります。これは、ガードレールが
|
||||
デバウンスされたトランスクリプトテキストに対して実行され、トリップワイヤーが発火した時点で一部の音声がすでにバッファリングされている可能性があるためです。
|
||||
|
||||
## SIP とテレフォニー
|
||||
|
||||
Python SDK には、[`OpenAIRealtimeSIPModel`][agents.realtime.openai_realtime.OpenAIRealtimeSIPModel] によるファーストクラスの SIP アタッチフローが含まれています。
|
||||
|
||||
Realtime Calls API 経由で着信があり、生成された `call_id` にエージェントセッションをアタッチしたい場合に使用します:
|
||||
|
||||
```python
|
||||
from agents.realtime import RealtimeRunner
|
||||
from agents.realtime.openai_realtime import OpenAIRealtimeSIPModel
|
||||
|
||||
runner = RealtimeRunner(starting_agent=agent, model=OpenAIRealtimeSIPModel())
|
||||
|
||||
async with await runner.run(
|
||||
model_config={
|
||||
"call_id": call_id_from_webhook,
|
||||
}
|
||||
) as session:
|
||||
async for event in session:
|
||||
...
|
||||
```
|
||||
|
||||
最初に通話を受け入れる必要があり、accept ペイロードをエージェントから導出されたセッション設定と一致させたい場合は、`OpenAIRealtimeSIPModel.build_initial_session_payload(...)` を使用してください。完全なフローは [`examples/realtime/twilio_sip/server.py`](https://github.com/openai/openai-agents-python/tree/main/examples/realtime/twilio_sip/server.py) に示されています。
|
||||
|
||||
## 低レベルアクセスとカスタムエンドポイント
|
||||
|
||||
`session.model` を通じて基盤となるトランスポートオブジェクトにアクセスできます。
|
||||
|
||||
次が必要な場合に使用します:
|
||||
|
||||
- `session.model.add_listener(...)` によるカスタムリスナー
|
||||
- `response.create` や `session.update` などの raw クライアントイベント
|
||||
- `model_config` を通じたカスタムの `url`、`headers`、`api_key` 処理
|
||||
- 既存のリアルタイム通話への `call_id` アタッチ
|
||||
|
||||
`RealtimeModelConfig` は次をサポートします:
|
||||
|
||||
- `api_key`
|
||||
- `url`
|
||||
- `headers`
|
||||
- `initial_model_settings`
|
||||
- `playback_tracker`
|
||||
- `call_id`
|
||||
|
||||
このリポジトリに同梱されている `call_id` の例は SIP です。より広範な Realtime API でも、一部のサーバー側制御フローで `call_id` が使用されますが、ここではそれらは Python のコード例としてパッケージ化されていません。
|
||||
|
||||
Azure OpenAI に接続する場合は、GA Realtime エンドポイント URL と明示的なヘッダーを渡してください。例:
|
||||
|
||||
```python
|
||||
session = await runner.run(
|
||||
model_config={
|
||||
"url": "wss://<your-resource>.openai.azure.com/openai/v1/realtime?model=<deployment-name>",
|
||||
"headers": {"api-key": "<your-azure-api-key>"},
|
||||
}
|
||||
)
|
||||
```
|
||||
|
||||
トークンベースの認証では、`headers` にベアラートークンを使用します:
|
||||
|
||||
```python
|
||||
session = await runner.run(
|
||||
model_config={
|
||||
"url": "wss://<your-resource>.openai.azure.com/openai/v1/realtime?model=<deployment-name>",
|
||||
"headers": {"authorization": f"Bearer {token}"},
|
||||
}
|
||||
)
|
||||
```
|
||||
|
||||
`headers` を渡した場合、SDK は `Authorization` を自動的に追加しません。リアルタイムエージェントでは、従来のベータパス(`/openai/realtime?api-version=...`)を避けてください。
|
||||
|
||||
## 参考資料
|
||||
|
||||
- [Realtime トランスポート](transport.md)
|
||||
- [クイックスタート](quickstart.md)
|
||||
- [OpenAI Realtime 会話](https://developers.openai.com/api/docs/guides/realtime-conversations/)
|
||||
- [OpenAI Realtime サーバー側制御](https://developers.openai.com/api/docs/guides/realtime-server-controls/)
|
||||
- [`examples/realtime`](https://github.com/openai/openai-agents-python/tree/main/examples/realtime)
|
||||
@@ -0,0 +1,162 @@
|
||||
---
|
||||
search:
|
||||
exclude: true
|
||||
---
|
||||
# クイックスタート
|
||||
|
||||
Python SDK のリアルタイムエージェントは、 WebSocket トランスポート経由の OpenAI Realtime API を基盤とする、サーバー側の低レイテンシエージェントです。
|
||||
|
||||
!!! warning "ベータ機能"
|
||||
|
||||
リアルタイムエージェントはベータ版です。実装を改善していく中で、破壊的変更が発生する可能性があります。
|
||||
|
||||
!!! note "Python SDK の範囲"
|
||||
|
||||
Python SDK はブラウザー向け WebRTC トランスポートを **提供しません** 。このページでは、サーバー側 WebSocket 上で Python が管理するリアルタイムセッションのみを扱います。この SDK は、サーバー側のオーケストレーション、ツール、承認、テレフォニー連携に使用してください。[リアルタイムトランスポート](transport.md) も参照してください。
|
||||
|
||||
## 前提条件
|
||||
|
||||
- Python 3.10 以上
|
||||
- OpenAI API キー
|
||||
- OpenAI Agents SDK に関する基本的な知識
|
||||
|
||||
## インストール
|
||||
|
||||
まだの場合は、 OpenAI Agents SDK をインストールしてください:
|
||||
|
||||
```bash
|
||||
pip install openai-agents
|
||||
```
|
||||
|
||||
## サーバー側リアルタイムセッションの作成
|
||||
|
||||
### 1. リアルタイムコンポーネントのインポート
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
|
||||
from agents.realtime import RealtimeAgent, RealtimeRunner
|
||||
```
|
||||
|
||||
### 2. 開始時のエージェントの定義
|
||||
|
||||
```python
|
||||
agent = RealtimeAgent(
|
||||
name="Assistant",
|
||||
instructions="You are a helpful voice assistant. Keep responses short and conversational.",
|
||||
)
|
||||
```
|
||||
|
||||
### 3. ランナーの設定
|
||||
|
||||
新しいコードでは、ネストされた `audio.input` / `audio.output` のセッション設定形式を推奨します。新しいリアルタイムエージェントでは、 `gpt-realtime-2` から始めてください。
|
||||
|
||||
```python
|
||||
runner = RealtimeRunner(
|
||||
starting_agent=agent,
|
||||
config={
|
||||
"model_settings": {
|
||||
"model_name": "gpt-realtime-2",
|
||||
"audio": {
|
||||
"input": {
|
||||
"format": "pcm16",
|
||||
"transcription": {"model": "gpt-4o-mini-transcribe"},
|
||||
"turn_detection": {
|
||||
"type": "semantic_vad",
|
||||
"interrupt_response": True,
|
||||
},
|
||||
},
|
||||
"output": {
|
||||
"format": "pcm16",
|
||||
"voice": "ash",
|
||||
},
|
||||
},
|
||||
}
|
||||
},
|
||||
)
|
||||
```
|
||||
|
||||
### 4. セッションの開始と入力の送信
|
||||
|
||||
`runner.run()` は `RealtimeSession` を返します。セッションコンテキストに入ると接続が開かれます。
|
||||
|
||||
```python
|
||||
async def main() -> None:
|
||||
session = await runner.run()
|
||||
|
||||
async with session:
|
||||
await session.send_message("Say hello in one short sentence.")
|
||||
|
||||
async for event in session:
|
||||
if event.type == "audio":
|
||||
# Forward or play event.audio.data.
|
||||
pass
|
||||
elif event.type == "history_added":
|
||||
print(event.item)
|
||||
elif event.type == "agent_end":
|
||||
# One assistant turn finished.
|
||||
break
|
||||
elif event.type == "error":
|
||||
print(f"Error: {event.error}")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
`session.send_message()` は、プレーンな文字列または構造化されたリアルタイムメッセージを受け付けます。未加工の音声チャンクには、 [`session.send_audio()`][agents.realtime.session.RealtimeSession.send_audio] を使用してください。
|
||||
|
||||
## このクイックスタートに含まれない内容
|
||||
|
||||
- マイクキャプチャとスピーカー再生のコード。リアルタイムのコード例については、 [`examples/realtime`](https://github.com/openai/openai-agents-python/tree/main/examples/realtime) を参照してください。
|
||||
- SIP / テレフォニーのアタッチフロー。[リアルタイムトランスポート](transport.md) と [SIP セクション](guide.md#sip-and-telephony) を参照してください。
|
||||
|
||||
## 主要設定
|
||||
|
||||
基本的なセッションが動作したら、多くの方が次に利用する設定は次のとおりです:
|
||||
|
||||
- `model_name`
|
||||
- `audio.input.format`, `audio.output.format`
|
||||
- `audio.input.transcription`
|
||||
- `audio.input.noise_reduction`
|
||||
- 自動ターン検出用の `audio.input.turn_detection`
|
||||
- `audio.output.voice`
|
||||
- `tool_choice`, `prompt`, `tracing`
|
||||
- `async_tool_calls`, `guardrails_settings.debounce_text_length`, `tool_error_formatter`
|
||||
|
||||
`input_audio_format`、`output_audio_format`、`input_audio_transcription`、`turn_detection` などの古いフラットなエイリアスも引き続き機能しますが、新しいコードではネストされた `audio` 設定を推奨します。
|
||||
|
||||
手動のターン制御では、 [リアルタイムエージェントガイド](guide.md#manual-response-control) で説明されているように、 raw な `session.update` / `input_audio_buffer.commit` / `response.create` フローを使用してください。
|
||||
|
||||
完全なスキーマについては、 [`RealtimeRunConfig`][agents.realtime.config.RealtimeRunConfig] と [`RealtimeSessionModelSettings`][agents.realtime.config.RealtimeSessionModelSettings] を参照してください。
|
||||
|
||||
## 接続オプション
|
||||
|
||||
環境で API キーを設定します:
|
||||
|
||||
```bash
|
||||
export OPENAI_API_KEY="your-api-key-here"
|
||||
```
|
||||
|
||||
または、セッションの開始時に直接渡します:
|
||||
|
||||
```python
|
||||
session = await runner.run(model_config={"api_key": "your-api-key"})
|
||||
```
|
||||
|
||||
`model_config` は以下もサポートしています:
|
||||
|
||||
- `url`: カスタム WebSocket エンドポイント
|
||||
- `headers`: カスタムリクエストヘッダー
|
||||
- `call_id`: 既存のリアルタイム通話にアタッチします。このリポジトリでは、ドキュメント化されているアタッチフローは SIP です。
|
||||
- `playback_tracker`: ユーザーが実際に聞いた音声の量を報告します
|
||||
|
||||
`headers` を明示的に渡す場合、 SDK は `Authorization` ヘッダーを **挿入しません** 。
|
||||
|
||||
Azure OpenAI に接続する場合は、 `model_config["url"]` に GA Realtime エンドポイント URL を指定し、ヘッダーを明示的に渡してください。リアルタイムエージェントでは、レガシーのベータパス (`/openai/realtime?api-version=...`) は避けてください。詳細は [リアルタイムエージェントガイド](guide.md#low-level-access-and-custom-endpoints) を参照してください。
|
||||
|
||||
## 次のステップ
|
||||
|
||||
- サーバー側 WebSocket と SIP のどちらを選ぶかについては、 [リアルタイムトランスポート](transport.md) をお読みください。
|
||||
- ライフサイクル、構造化入力、承認、ハンドオフ、ガードレール、低レベル制御については、 [リアルタイムエージェントガイド](guide.md) をお読みください。
|
||||
- [`examples/realtime`](https://github.com/openai/openai-agents-python/tree/main/examples/realtime) のコード例を参照してください。
|
||||
@@ -0,0 +1,76 @@
|
||||
---
|
||||
search:
|
||||
exclude: true
|
||||
---
|
||||
# Realtime トランスポート
|
||||
|
||||
このページは、realtime エージェントを Python アプリケーションにどのように組み込むかを判断するために使用してください。
|
||||
|
||||
!!! note "Python SDK の境界"
|
||||
|
||||
Python SDK には、ブラウザー WebRTC トランスポートは含まれて **いません**。このページは、Python SDK のトランスポート選択肢であるサーバー側 WebSocket と SIP アタッチフローのみを扱います。ブラウザー WebRTC は別のプラットフォームトピックであり、公式の [WebRTC による Realtime API](https://developers.openai.com/api/docs/guides/realtime-webrtc/) ガイドに記載されています。
|
||||
|
||||
## 判断ガイド
|
||||
|
||||
| 目的 | はじめに | 理由 |
|
||||
| --- | --- | --- |
|
||||
| サーバー管理の realtime アプリを構築する | [クイックスタート](quickstart.md) | デフォルトの Python パスは、`RealtimeRunner` によって管理されるサーバー側 WebSocket セッションです。 |
|
||||
| 選択すべきトランスポートとデプロイ形態を理解する | このページ | トランスポートやデプロイ形態を決定する前に使用してください。 |
|
||||
| エージェントを電話または SIP 通話にアタッチする | [Realtime ガイド](guide.md) と [`examples/realtime/twilio_sip`](https://github.com/openai/openai-agents-python/tree/main/examples/realtime/twilio_sip) | このリポジトリには、`call_id` によって駆動される SIP アタッチフローが含まれています。 |
|
||||
|
||||
## デフォルトの Python パスであるサーバー側 WebSocket
|
||||
|
||||
カスタム `RealtimeModel` を渡さない限り、`RealtimeRunner` は `OpenAIRealtimeWebSocketModel` を使用します。
|
||||
|
||||
つまり、標準的な Python トポロジーは次のようになります。
|
||||
|
||||
1. Python サービスが `RealtimeRunner` を作成します。
|
||||
2. `await runner.run()` が `RealtimeSession` を返します。
|
||||
3. セッションに入り、テキスト、構造化メッセージ、または音声を送信します。
|
||||
4. `RealtimeSessionEvent` 項目を消費し、音声またはトランスクリプトをアプリケーションに転送します。
|
||||
|
||||
これは、コアデモアプリ、CLI の例、Twilio Media Streams の例で使用されているトポロジーです。
|
||||
|
||||
- [`examples/realtime/app`](https://github.com/openai/openai-agents-python/tree/main/examples/realtime/app)
|
||||
- [`examples/realtime/cli`](https://github.com/openai/openai-agents-python/tree/main/examples/realtime/cli)
|
||||
- [`examples/realtime/twilio`](https://github.com/openai/openai-agents-python/tree/main/examples/realtime/twilio)
|
||||
|
||||
サーバーが音声パイプライン、ツール実行、承認フロー、履歴処理を管理する場合は、このパスを使用してください。
|
||||
|
||||
## テレフォニー向けパスとしての SIP アタッチ
|
||||
|
||||
このリポジトリで説明されているテレフォニーフローでは、Python SDK は `call_id` を介して既存の realtime 通話にアタッチします。
|
||||
|
||||
このトポロジーは次のようになります。
|
||||
|
||||
1. OpenAI が `realtime.call.incoming` などの webhook をサービスに送信します。
|
||||
2. サービスが Realtime Calls API を通じて通話を受け入れます。
|
||||
3. Python サービスが `RealtimeRunner(..., model=OpenAIRealtimeSIPModel())` を開始します。
|
||||
4. セッションは `model_config={"call_id": ...}` で接続し、その後は他の realtime セッションと同様にイベントを処理します。
|
||||
|
||||
これは [`examples/realtime/twilio_sip`](https://github.com/openai/openai-agents-python/tree/main/examples/realtime/twilio_sip) に示されているトポロジーです。
|
||||
|
||||
より広範な Realtime API でも、一部のサーバー側制御パターンで `call_id` を使用しますが、このリポジトリに含まれるアタッチ例は SIP です。
|
||||
|
||||
## この SDK の範囲外であるブラウザー WebRTC
|
||||
|
||||
アプリの主なクライアントが Realtime WebRTC を使用するブラウザーである場合:
|
||||
|
||||
- このリポジトリの Python SDK ドキュメントの範囲外として扱ってください。
|
||||
- クライアント側のフローとイベントモデルについては、公式の [WebRTC による Realtime API](https://developers.openai.com/api/docs/guides/realtime-webrtc/) および [Realtime conversations](https://developers.openai.com/api/docs/guides/realtime-conversations/) ドキュメントを使用してください。
|
||||
- ブラウザー WebRTC クライアントの上にサイドバンドのサーバー接続が必要な場合は、公式の [Realtime server-side controls](https://developers.openai.com/api/docs/guides/realtime-server-controls/) ガイドを使用してください。
|
||||
- このリポジトリが、ブラウザー側の `RTCPeerConnection` 抽象化や、すぐに使えるブラウザー WebRTC サンプルを提供することは期待しないでください。
|
||||
|
||||
このリポジトリには、現在、ブラウザー WebRTC と Python サイドバンドを組み合わせた例も含まれていません。
|
||||
|
||||
## カスタムエンドポイントとアタッチポイント
|
||||
|
||||
[`RealtimeModelConfig`][agents.realtime.model.RealtimeModelConfig] のトランスポート設定サーフェスを使用すると、デフォルトのパスを調整できます。
|
||||
|
||||
- `url`: WebSocket エンドポイントを上書きします
|
||||
- `headers`: Azure 認証ヘッダーなどの明示的なヘッダーを指定します
|
||||
- `api_key`: API キーを直接、またはコールバック経由で渡します
|
||||
- `call_id`: 既存の realtime 通話にアタッチします。このリポジトリで記載されている例は SIP です。
|
||||
- `playback_tracker`: 割り込み処理のために実際の再生進捗を報告します
|
||||
|
||||
トポロジーを選択した後の詳細なライフサイクルと機能サーフェスについては、[Realtime エージェントガイド](guide.md) を参照してください。
|
||||
@@ -0,0 +1,178 @@
|
||||
---
|
||||
search:
|
||||
exclude: true
|
||||
---
|
||||
# リリースプロセス / 変更履歴
|
||||
|
||||
このプロジェクトは、形式 `0.Y.Z` を使用するセマンティックバージョニングを少し修正したものに従います。先頭の `0` は、SDK がまだ急速に進化中であることを示します。各構成要素は次のように増加させます:
|
||||
|
||||
## マイナー (`Y`) バージョン
|
||||
|
||||
ベータとしてマークされていない公開インターフェイスに対する **破壊的変更** の場合、マイナーバージョン `Y` を増やします。たとえば、`0.0.x` から `0.1.x` への移行には破壊的変更が含まれる可能性があります。
|
||||
|
||||
破壊的変更を避けたい場合は、プロジェクトで `0.0.x` バージョンに固定することをおすすめします。
|
||||
|
||||
## パッチ (`Z`) バージョン
|
||||
|
||||
非破壊的変更の場合は `Z` を増やします:
|
||||
|
||||
- バグ修正
|
||||
- 新機能
|
||||
- 非公開インターフェイスの変更
|
||||
- ベータ機能の更新
|
||||
|
||||
## 破壊的変更の変更履歴
|
||||
|
||||
### 0.17.0
|
||||
|
||||
このバージョンでは、サンドボックスのローカルソースの実体化において、ソースパスが `Manifest.extra_path_grants` によってカバーされていない限り、`LocalFile.src` と `LocalDir.src` は実体化時の `base_dir` 内に収められます。`base_dir` は、マニフェストが適用される時点での SDK プロセスの現在の作業ディレクトリです。相対ローカルソースはそのディレクトリから解決され、絶対ローカルソースはすでにそのディレクトリ内にあるか、明示的な許可の配下にある必要があります。これによりローカルアーティファクトの境界に関する問題は解消されますが、そのベースディレクトリ外から信頼済みホストのファイルまたはディレクトリをサンドボックスワークスペースへ意図的にコピーするアプリケーションに影響する可能性があります。
|
||||
|
||||
移行するには、信頼済みホストルートをマニフェストレベルで `SandboxPathGrant` により許可してください。サンドボックスがそれらのファイルを読み取るだけでよい場合は、読み取り専用にすることをおすすめします:
|
||||
|
||||
```python
|
||||
from pathlib import Path
|
||||
|
||||
from agents.sandbox import Manifest, SandboxPathGrant
|
||||
from agents.sandbox.entries import Dir, LocalDir
|
||||
|
||||
# This is an absolute host path outside the SDK process base_dir.
|
||||
TRUSTED_DOCS_ROOT = Path("/opt/my-app/docs")
|
||||
|
||||
manifest = Manifest(
|
||||
extra_path_grants=(
|
||||
# This host root is outside the SDK process base_dir, so the manifest must grant it.
|
||||
SandboxPathGrant(path=str(TRUSTED_DOCS_ROOT), read_only=True),
|
||||
),
|
||||
entries={
|
||||
# No grant is needed for local sources that stay under the SDK process base_dir.
|
||||
"fixtures": LocalDir(src=Path("fixtures"), description="Local test fixtures."),
|
||||
# This entry reads from the granted host root and copies it into the sandbox workspace.
|
||||
"docs": LocalDir(src=TRUSTED_DOCS_ROOT, description="Trusted local documents."),
|
||||
# Dir creates a sandbox workspace directory; it does not read from the host filesystem.
|
||||
"output": Dir(description="Generated artifacts."),
|
||||
},
|
||||
)
|
||||
```
|
||||
|
||||
`extra_path_grants` は信頼済みアプリケーション設定として扱ってください。アプリケーションがそれらのホストパスをすでに承認していない限り、モデル出力やその他の信頼できないマニフェスト入力から許可を設定しないでください。
|
||||
|
||||
### 0.16.0
|
||||
|
||||
このバージョンでは、SDK のデフォルトモデルが `gpt-4.1` ではなく `gpt-5.4-mini` になりました。これは、モデルを明示的に設定していないエージェントと実行に影響します。新しいデフォルトは GPT-5 モデルであるため、暗黙的なデフォルトモデル設定には `reasoning.effort="none"` や `verbosity="low"` などの GPT-5 デフォルトが含まれるようになりました。
|
||||
|
||||
以前のデフォルトモデルの挙動を維持する必要がある場合は、エージェントまたは実行設定でモデルを明示的に設定するか、`OPENAI_DEFAULT_MODEL` 環境変数を設定してください:
|
||||
|
||||
```python
|
||||
agent = Agent(name="Assistant", model="gpt-4.1")
|
||||
```
|
||||
|
||||
主な変更点:
|
||||
|
||||
- `Runner.run`、`Runner.run_sync`、`Runner.run_streamed` は、ターン制限を無効にするために `max_turns=None` を受け取れるようになりました。
|
||||
- サンドボックスワークスペースのハイドレーションは、ローカル、Docker、およびプロバイダーがバックするサンドボックス実装全体で、絶対シンボリックリンクターゲットを含む、アーカイブルートの外部を指すシンボリックリンクを含む tar アーカイブを拒否するようになりました。
|
||||
|
||||
### 0.15.0
|
||||
|
||||
このバージョンでは、モデルの拒否応答は、空のテキスト出力として扱われたり、structured outputs の場合に実行ループが `MaxTurnsExceeded` まで再試行したりするのではなく、`ModelRefusalError` として明示的に表面化されるようになりました。
|
||||
|
||||
これは、以前に拒否のみのモデル応答が `final_output == ""` で完了することを期待していたコードに影響します。例外を発生させずに拒否を処理するには、`model_refusal` 実行エラーハンドラーを提供してください:
|
||||
|
||||
```python
|
||||
result = Runner.run_sync(
|
||||
agent,
|
||||
input,
|
||||
error_handlers={"model_refusal": lambda data: data.error.refusal},
|
||||
)
|
||||
```
|
||||
|
||||
structured-output エージェントの場合、ハンドラーはエージェントの出力スキーマに一致する値を返すことができ、SDK は他の実行エラーハンドラーの最終出力と同様に検証します。
|
||||
|
||||
### 0.14.0
|
||||
|
||||
このマイナーリリースでは、破壊的変更は **導入しません** が、主要な新しいベータ機能領域である Sandbox エージェントに加え、ローカル、コンテナ化、ホスト環境全体でそれらを使用するために必要なランタイム、バックエンド、ドキュメントのサポートが追加されています。
|
||||
|
||||
主な変更点:
|
||||
|
||||
- `SandboxAgent`、`Manifest`、`SandboxRunConfig` を中心とした新しいベータのサンドボックスランタイムサーフェスを追加しました。これにより、エージェントはファイル、ディレクトリ、Git リポジトリ、マウント、スナップショット、再開サポートを備えた永続的な隔離ワークスペース内で動作できます。
|
||||
- `UnixLocalSandboxClient` と `DockerSandboxClient` によるローカルおよびコンテナ化開発向けのサンドボックス実行バックエンドに加え、任意の extras を通じた Blaxel、Cloudflare、Daytona、E2B、Modal、Runloop、Vercel 向けのホスト型プロバイダー連携を追加しました。
|
||||
- 将来の実行が以前の実行から得た知見を再利用できるように、サンドボックスメモリサポートを追加しました。段階的開示、複数ターンのグループ化、設定可能な隔離境界、および S3 バックのワークフローを含む永続化メモリのコード例が含まれます。
|
||||
- ローカルおよび合成ワークスペースエントリー、S3/R2/GCS/Azure Blob Storage/S3 Files 向けのリモートストレージマウント、移植可能なスナップショット、`RunState`、`SandboxSessionState`、または保存済みスナップショットによる再開フローを含む、より広範なワークスペースと再開モデルを追加しました。
|
||||
- `examples/sandbox/` 配下に、充実したサンドボックスのコード例とチュートリアルを追加しました。スキル、ハンドオフ、メモリを用いたコーディングタスク、プロバイダー固有のセットアップ、コードレビュー、データルーム QA、Web サイトのクローン作成などのエンドツーエンドのワークフローを扱います。
|
||||
- サンドボックス対応のセッション準備、ケイパビリティバインディング、状態のシリアライズ、統合トレーシング、プロンプトキャッシュキーのデフォルト、より安全な機密 MCP 出力のマスキングにより、コアランタイムとトレーシングスタックを拡張しました。
|
||||
|
||||
### 0.13.0
|
||||
|
||||
このマイナーリリースでは、破壊的変更は **導入しません** が、注目すべき Realtime のデフォルト更新に加え、新しい MCP 機能とランタイム安定性の修正が含まれます。
|
||||
|
||||
主な変更点:
|
||||
|
||||
- デフォルトの WebSocket Realtime モデルは `gpt-realtime-1.5` になりました。そのため、新しい Realtime エージェントのセットアップでは、追加設定なしで新しいモデルが使用されます。
|
||||
- `MCPServer` は `list_resources()`、`list_resource_templates()`、`read_resource()` を公開するようになりました。また、`MCPServerStreamableHttp` は `session_id` を公開するようになったため、ストリーム可能な HTTP セッションを再接続やステートレスワーカーをまたいで再開できます。
|
||||
- Chat Completions 連携は、`should_replay_reasoning_content` によって推論コンテンツの再生をオプトインできるようになりました。これにより、LiteLLM/DeepSeek などのアダプターで、プロバイダー固有の推論 / ツール呼び出しの連続性が向上します。
|
||||
- `SQLAlchemySession` における初回書き込みの同時実行、推論の除去後に孤立した assistant メッセージ ID を持つ圧縮リクエスト、`remove_all_tools()` が MCP/reasoning 項目を残す問題、関数ツールバッチ実行器の競合など、複数のランタイムおよびセッションのエッジケースを修正しました。
|
||||
|
||||
### 0.12.0
|
||||
|
||||
このマイナーリリースでは、破壊的変更は **導入しません**。主要な機能追加については、[リリースノート](https://github.com/openai/openai-agents-python/releases/tag/v0.12.0)を確認してください。
|
||||
|
||||
### 0.11.0
|
||||
|
||||
このマイナーリリースでは、破壊的変更は **導入しません**。主要な機能追加については、[リリースノート](https://github.com/openai/openai-agents-python/releases/tag/v0.11.0)を確認してください。
|
||||
|
||||
### 0.10.0
|
||||
|
||||
このマイナーリリースでは、破壊的変更は **導入しません** が、OpenAI Responses ユーザー向けの重要な新機能領域である Responses API の WebSocket トランスポートサポートが含まれます。
|
||||
|
||||
主な変更点:
|
||||
|
||||
- OpenAI Responses モデル向けの WebSocket トランスポートサポートを追加しました(オプトインです。HTTP は引き続きデフォルトのトランスポートです)。
|
||||
- 複数ターンの実行全体で共有の WebSocket 対応プロバイダーと `RunConfig` を再利用するための `responses_websocket_session()` ヘルパー / `ResponsesWebSocketSession` を追加しました。
|
||||
- ストリーミング、ツール、承認、フォローアップターンを扱う新しい WebSocket ストリーミングコード例(`examples/basic/stream_ws.py`)を追加しました。
|
||||
|
||||
### 0.9.0
|
||||
|
||||
このバージョンでは、このメジャーバージョンが 3 か月前に EOL に達したため、Python 3.9 はサポートされなくなりました。より新しいランタイムバージョンへアップグレードしてください。
|
||||
|
||||
さらに、`Agent#as_tool()` メソッドから返される値の型ヒントが、`Tool` から `FunctionTool` へ狭められました。この変更は通常、破壊的な問題を引き起こすことはありませんが、コードがより広い Union 型に依存している場合は、側でいくつか調整が必要になる可能性があります。
|
||||
|
||||
### 0.8.0
|
||||
|
||||
このバージョンでは、2 つのランタイム挙動の変更により、移行作業が必要になる可能性があります:
|
||||
|
||||
- **同期** Python 呼び出し可能オブジェクトをラップする関数ツールは、イベントループスレッドで実行されるのではなく、`asyncio.to_thread(...)` を介してワーカースレッドで実行されるようになりました。ツールのロジックがスレッドローカル状態やスレッドアフィンなリソースに依存している場合は、async ツール実装へ移行するか、ツールコード内でスレッドアフィニティを明示してください。
|
||||
- ローカル MCP ツールの失敗処理は設定可能になり、デフォルトの挙動では実行全体を失敗させる代わりに、モデルから見えるエラー出力を返す場合があります。フェイルファストのセマンティクスに依存している場合は、`mcp_config={"failure_error_function": None}` を設定してください。サーバーレベルの `failure_error_function` 値はエージェントレベルの設定を上書きするため、明示的なハンドラーを持つ各ローカル MCP サーバーで `failure_error_function=None` を設定してください。
|
||||
|
||||
### 0.7.0
|
||||
|
||||
このバージョンでは、既存のアプリケーションに影響する可能性がある挙動の変更がいくつかありました:
|
||||
|
||||
- ネストされたハンドオフ履歴は **オプトイン** になりました(デフォルトでは無効)。v0.6.x のデフォルトのネスト挙動に依存していた場合は、`RunConfig(nest_handoff_history=True)` を明示的に設定してください。
|
||||
- `gpt-5.1` / `gpt-5.2` のデフォルトの `reasoning.effort` は、`"none"` に変更されました(SDK デフォルトで設定されていた以前のデフォルト `"low"` からの変更です)。プロンプトや品質 / コストのプロファイルが `"low"` に依存していた場合は、`model_settings` で明示的に設定してください。
|
||||
|
||||
### 0.6.0
|
||||
|
||||
このバージョンでは、デフォルトのハンドオフ履歴は、生のユーザー / アシスタントターンを公開するのではなく、単一の assistant メッセージにまとめられるようになり、後続のエージェントに簡潔で予測可能な要約を提供します
|
||||
- 既存の単一メッセージのハンドオフトランスクリプトは、デフォルトで `<CONVERSATION HISTORY>` ブロックの前に "For context, here is the conversation so far between the user and the previous agent:" で始まるようになったため、後続のエージェントは明確なラベル付きの要約を受け取れます
|
||||
|
||||
### 0.5.0
|
||||
|
||||
このバージョンでは、目に見える破壊的変更は導入されませんが、新機能と内部のいくつかの重要な更新が含まれます:
|
||||
|
||||
- `RealtimeRunner` が [SIP プロトコル接続](https://platform.openai.com/docs/guides/realtime-sip)を処理するためのサポートを追加しました
|
||||
- Python 3.14 互換性のために、`Runner#run_sync` の内部ロジックを大幅に改訂しました
|
||||
|
||||
### 0.4.0
|
||||
|
||||
このバージョンでは、[openai](https://pypi.org/project/openai/) パッケージの v1.x バージョンはサポートされなくなりました。この SDK とともに openai v2.x を使用してください。
|
||||
|
||||
### 0.3.0
|
||||
|
||||
このバージョンでは、Realtime API サポートは gpt-realtime モデルとその API インターフェイス(GA バージョン)へ移行します。
|
||||
|
||||
### 0.2.0
|
||||
|
||||
このバージョンでは、以前は `Agent` を引数として受け取っていたいくつかの箇所が、代わりに `AgentBase` を引数として受け取るようになりました。たとえば、MCP サーバーの `list_tools()` 呼び出しです。これは純粋に型付け上の変更であり、引き続き `Agent` オブジェクトを受け取ります。更新するには、`Agent` を `AgentBase` に置き換えて型エラーを修正するだけです。
|
||||
|
||||
### 0.1.0
|
||||
|
||||
このバージョンでは、[`MCPServer.list_tools()`][agents.mcp.server.MCPServer] に `run_context` と `agent` という 2 つの新しいパラメーターが追加されました。`MCPServer` をサブクラス化しているすべてのクラスに、これらのパラメーターを追加する必要があります。
|
||||
@@ -0,0 +1,24 @@
|
||||
---
|
||||
search:
|
||||
exclude: true
|
||||
---
|
||||
# REPL ユーティリティ
|
||||
|
||||
SDK は、ターミナルでエージェントの動作を直接すばやく対話的にテストするための `run_demo_loop` を提供します。
|
||||
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
from agents import Agent, run_demo_loop
|
||||
|
||||
async def main() -> None:
|
||||
agent = Agent(name="Assistant", instructions="You are a helpful assistant.")
|
||||
await run_demo_loop(agent)
|
||||
|
||||
if __name__ == "__main__":
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
`run_demo_loop` はループ内でユーザー入力を求め、ターン間の会話履歴を保持します。デフォルトでは、生成されるモデル出力をストリーミングします。上記の例を実行すると、run_demo_loop は対話型チャットセッションを開始します。入力を継続的に求め、ターン間の会話履歴全体を記憶し(そのためエージェントはこれまでに話し合われた内容を把握できます)、エージェントの応答が生成されるとリアルタイムで自動的にストリーミングします。
|
||||
|
||||
このチャットセッションを終了するには、単に `quit` または `exit` と入力して Enter キーを押すか、`Ctrl-D` キーボードショートカットを使用します。
|
||||
+140
-27
@@ -1,52 +1,165 @@
|
||||
# 結果
|
||||
---
|
||||
search:
|
||||
exclude: true
|
||||
---
|
||||
# 実行結果
|
||||
|
||||
`Runner.run` メソッドを呼び出すと、以下のいずれかが返されます:
|
||||
`Runner.run` メソッドを呼び出すと、次の 2 つの実行結果型のいずれかを受け取ります。
|
||||
|
||||
- `run` または `run_sync` を呼び出した場合は、[`RunResult`][agents.result.RunResult]
|
||||
- `run_streamed` を呼び出した場合は、[`RunResultStreaming`][agents.result.RunResultStreaming]
|
||||
- `Runner.run(...)` または `Runner.run_sync(...)` からの [`RunResult`][agents.result.RunResult]
|
||||
- `Runner.run_streamed(...)` からの [`RunResultStreaming`][agents.result.RunResultStreaming]
|
||||
|
||||
これらはどちらも [`RunResultBase`][agents.result.RunResultBase] を継承しており、ほとんどの有用な情報はここに含まれています。
|
||||
どちらも [`RunResultBase`][agents.result.RunResultBase] を継承しており、`final_output`、`new_items`、`last_agent`、`raw_responses`、`to_state()` などの共通の実行結果サーフェスを公開します。
|
||||
|
||||
`RunResultStreaming` は、[`stream_events()`][agents.result.RunResultStreaming.stream_events]、[`current_agent`][agents.result.RunResultStreaming.current_agent]、[`is_complete`][agents.result.RunResultStreaming.is_complete]、[`cancel(...)`][agents.result.RunResultStreaming.cancel] など、ストリーミング固有の制御機能を追加します。
|
||||
|
||||
## 適切な実行結果サーフェスの選択
|
||||
|
||||
ほとんどのアプリケーションでは、いくつかの実行結果プロパティまたはヘルパーだけで十分です。
|
||||
|
||||
| 必要なもの... | 使用するもの |
|
||||
| --- | --- |
|
||||
| ユーザーに表示する最終回答 | `final_output` |
|
||||
| 完全なローカルトランスクリプトを含む、リプレイ可能な次ターン入力リスト | `to_input_list()` |
|
||||
| エージェント、ツール、ハンドオフ、承認メタデータを含む詳細な実行項目 | `new_items` |
|
||||
| 通常、次のユーザーターンを処理すべきエージェント | `last_agent` |
|
||||
| `previous_response_id` による OpenAI Responses API チェーン | `last_response_id` |
|
||||
| 保留中の承認と再開可能なスナップショット | `interruptions` and `to_state()` |
|
||||
| 現在のネストされた `Agent.as_tool()` 呼び出しに関するメタデータ | `agent_tool_invocation` |
|
||||
| raw モデル呼び出しまたはガードレール診断 | `raw_responses` and the guardrail result arrays |
|
||||
|
||||
## 最終出力
|
||||
|
||||
[`final_output`][agents.result.RunResultBase.final_output] プロパティには、最後に実行されたエージェントの最終出力が格納されます。これは以下のいずれかです:
|
||||
[`final_output`][agents.result.RunResultBase.final_output] プロパティには、最後に実行されたエージェントの最終出力が含まれます。これは次のいずれかです。
|
||||
|
||||
- 最後のエージェントに `output_type` が定義されていない場合は `str`
|
||||
- エージェントに `output_type` が定義されている場合は、`last_agent.output_type` 型のオブジェクト
|
||||
- 最後のエージェントに `output_type` が定義されていなかった場合は `str`
|
||||
- 最後のエージェントに出力型が定義されていた場合は `last_agent.output_type` 型のオブジェクト
|
||||
- 承認中断で一時停止した場合など、最終出力が生成される前に実行が停止した場合は `None`
|
||||
|
||||
!!! note
|
||||
|
||||
`final_output` の型は `Any` です。ハンドオフのため、静的に型を決定することはできません。ハンドオフが発生した場合、どのエージェントが最後になるか分からないため、出力型の集合を静的に知ることができません。
|
||||
`final_output` は `Any` として型付けされています。ハンドオフによって、どのエージェントが実行を終了するかが変わる可能性があるため、SDK は考えられる出力型の全体集合を静的に把握できません。
|
||||
|
||||
## 次のターンへの入力
|
||||
ストリーミングモードでは、ストリームの処理が完了するまで `final_output` は `None` のままです。イベントごとのフローについては、[ストリーミング](streaming.md)を参照してください。
|
||||
|
||||
[`result.to_input_list()`][agents.result.RunResultBase.to_input_list] を使うことで、実行結果を入力リストに変換できます。これは、あなたが提供した元の入力と、エージェント実行中に生成されたアイテムを連結したものです。これにより、あるエージェント実行の出力を別の実行に渡したり、ループで実行して毎回新しいユーザー入力を追加したりするのが簡単になります。
|
||||
## 入力、次ターン履歴、新規項目
|
||||
|
||||
## 最後のエージェント
|
||||
これらのサーフェスは、それぞれ異なる問いに対応します。
|
||||
|
||||
[`last_agent`][agents.result.RunResultBase.last_agent] プロパティには、最後に実行されたエージェントが格納されます。アプリケーションによっては、次回ユーザーが何か入力した際にこれが役立つことがよくあります。たとえば、フロントラインのトリアージエージェントが言語特化エージェントにハンドオフする場合、最後のエージェントを保存しておき、次回ユーザーがエージェントにメッセージを送る際に再利用できます。
|
||||
| プロパティまたはヘルパー | 含まれる内容 | 最適な用途 |
|
||||
| --- | --- | --- |
|
||||
| [`input`][agents.result.RunResultBase.input] | この実行セグメントの基本入力です。ハンドオフ入力フィルターが履歴を書き換えた場合、実行が継続されたフィルター済み入力がここに反映されます。 | この実行が実際に入力として使用した内容の監査 |
|
||||
| [`to_input_list()`][agents.result.RunResultBase.to_input_list] | 実行を入力項目として見たビューです。デフォルトの `mode="preserve_all"` は、`new_items` から変換された完全な履歴を保持します。`mode="normalized"` は、ハンドオフフィルタリングによってモデル履歴が書き換えられた場合に、正規の継続入力を優先します。 | 手動のチャットループ、クライアント管理の会話状態、プレーンな項目履歴の確認 |
|
||||
| [`new_items`][agents.result.RunResultBase.new_items] | エージェント、ツール、ハンドオフ、承認メタデータを含む詳細な [`RunItem`][agents.items.RunItem] ラッパーです。 | ログ、UI、監査、デバッグ |
|
||||
| [`raw_responses`][agents.result.RunResultBase.raw_responses] | 実行内の各モデル呼び出しからの raw [`ModelResponse`][agents.items.ModelResponse] オブジェクトです。 | プロバイダーレベルの診断または raw レスポンスの確認 |
|
||||
|
||||
## 新しいアイテム
|
||||
実際には、次のように使い分けます。
|
||||
|
||||
[`new_items`][agents.result.RunResultBase.new_items] プロパティには、実行中に生成された新しいアイテムが格納されます。これらのアイテムは [`RunItem`][agents.items.RunItem] です。RunItem は LLM によって生成された raw アイテムをラップします。
|
||||
- 実行のプレーンな入力項目ビューが必要な場合は、`to_input_list()` を使用します。
|
||||
- ハンドオフフィルタリングまたはネストされたハンドオフ履歴の書き換え後に、次の `Runner.run(..., input=...)` 呼び出しに渡す正規のローカル入力が必要な場合は、`to_input_list(mode="normalized")` を使用します。
|
||||
- SDK に履歴の読み込みと保存を任せたい場合は、[`session=...`](sessions/index.md) を使用します。
|
||||
- `conversation_id` または `previous_response_id` を使って OpenAI のサーバー管理状態を使用している場合、通常は `to_input_list()` を再送信する代わりに、新しいユーザー入力のみを渡して保存済み ID を再利用します。
|
||||
- ログ、UI、監査向けに完全な変換済み履歴が必要な場合は、デフォルトの `to_input_list()` モードまたは `new_items` を使用します。
|
||||
|
||||
- [`MessageOutputItem`][agents.items.MessageOutputItem] は LLM からのメッセージを示します。raw アイテムは生成されたメッセージです。
|
||||
- [`HandoffCallItem`][agents.items.HandoffCallItem] は LLM がハンドオフツールを呼び出したことを示します。raw アイテムは LLM からのツールコールアイテムです。
|
||||
- [`HandoffOutputItem`][agents.items.HandoffOutputItem] はハンドオフが発生したことを示します。raw アイテムはハンドオフツールコールへのツールレスポンスです。また、アイテムからソース/ターゲットエージェントにもアクセスできます。
|
||||
- [`ToolCallItem`][agents.items.ToolCallItem] は LLM がツールを呼び出したことを示します。
|
||||
- [`ToolCallOutputItem`][agents.items.ToolCallOutputItem] はツールが呼び出されたことを示します。raw アイテムはツールレスポンスです。また、アイテムからツール出力にもアクセスできます。
|
||||
- [`ReasoningItem`][agents.items.ReasoningItem] は LLM からの推論アイテムを示します。raw アイテムは生成された推論です。
|
||||
JavaScript SDK とは異なり、Python ではモデル形式の差分のみを表す個別の `output` プロパティは公開されません。SDK メタデータが必要な場合は `new_items` を使用し、raw モデルペイロードが必要な場合は `raw_responses` を確認してください。
|
||||
|
||||
## その他の情報
|
||||
コンピュータツールのリプレイは、raw Responses ペイロードの形状に従います。プレビューモデルの `computer_call` 項目は単一の `action` を保持しますが、`gpt-5.5` のコンピュータ呼び出しではバッチ化された `actions[]` を保持できます。[`to_input_list()`][agents.result.RunResultBase.to_input_list] と [`RunState`][agents.run_state.RunState] は、モデルが生成した形状をそのまま保持するため、手動リプレイ、一時停止/再開フロー、保存済みトランスクリプトは、プレビュー版と GA 版の両方のコンピュータツール呼び出しで引き続き機能します。ローカル実行結果は引き続き `new_items` 内の `computer_call_output` 項目として表示されます。
|
||||
|
||||
### ガードレール結果
|
||||
### 新規項目
|
||||
|
||||
[`input_guardrail_results`][agents.result.RunResultBase.input_guardrail_results] および [`output_guardrail_results`][agents.result.RunResultBase.output_guardrail_results] プロパティには、ガードレールの結果(存在する場合)が格納されます。ガードレール結果には、ログや保存に役立つ有用な情報が含まれることがあるため、これらを利用できるようにしています。
|
||||
[`new_items`][agents.result.RunResultBase.new_items] は、実行中に何が起きたかを最も詳細に確認できるビューです。一般的な項目型は次のとおりです。
|
||||
|
||||
- アシスタントメッセージ用の [`MessageOutputItem`][agents.items.MessageOutputItem]
|
||||
- 推論項目用の [`ReasoningItem`][agents.items.ReasoningItem]
|
||||
- Responses のツール検索リクエストと、ロードされたツール検索の実行結果用の [`ToolSearchCallItem`][agents.items.ToolSearchCallItem] および [`ToolSearchOutputItem`][agents.items.ToolSearchOutputItem]
|
||||
- ツール呼び出しとその実行結果用の [`ToolCallItem`][agents.items.ToolCallItem] および [`ToolCallOutputItem`][agents.items.ToolCallOutputItem]
|
||||
- 承認待ちで一時停止したツール呼び出し用の [`ToolApprovalItem`][agents.items.ToolApprovalItem]
|
||||
- ハンドオフリクエストと完了済みの引き継ぎ用の [`HandoffCallItem`][agents.items.HandoffCallItem] および [`HandoffOutputItem`][agents.items.HandoffOutputItem]
|
||||
|
||||
エージェントの関連付け、ツール出力、ハンドオフの境界、承認の境界が必要な場合は、常に `to_input_list()` よりも `new_items` を選択してください。
|
||||
|
||||
ホスト型ツール検索を使用する場合、モデルが生成した検索リクエストを確認するには `ToolSearchCallItem.raw_item` を確認し、そのターンでどの名前空間、関数、またはホスト型 MCP サーバーがロードされたかを確認するには `ToolSearchOutputItem.raw_item` を確認してください。
|
||||
|
||||
## 会話の継続または再開
|
||||
|
||||
### 次ターンのエージェント
|
||||
|
||||
[`last_agent`][agents.result.RunResultBase.last_agent] には、最後に実行されたエージェントが含まれます。これは多くの場合、ハンドオフ後の次のユーザーターンで再利用するのに最適なエージェントです。
|
||||
|
||||
ストリーミングモードでは、[`RunResultStreaming.current_agent`][agents.result.RunResultStreaming.current_agent] が実行の進行に合わせて更新されるため、ストリームが終了する前にハンドオフを観察できます。
|
||||
|
||||
### 中断と実行状態
|
||||
|
||||
ツールに承認が必要な場合、保留中の承認は [`RunResult.interruptions`][agents.result.RunResult.interruptions] または [`RunResultStreaming.interruptions`][agents.result.RunResultStreaming.interruptions] で公開されます。これには、直接呼び出されたツール、ハンドオフ後に到達したツール、またはネストされた [`Agent.as_tool()`][agents.agent.Agent.as_tool] 実行によって発生した承認が含まれることがあります。
|
||||
|
||||
[`to_state()`][agents.result.RunResult.to_state] を呼び出して、再開可能な [`RunState`][agents.run_state.RunState] を取得し、保留中の項目を承認または拒否してから、`Runner.run(...)` または `Runner.run_streamed(...)` で再開します。
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner
|
||||
|
||||
agent = Agent(name="Assistant", instructions="Use tools when needed.")
|
||||
result = await Runner.run(agent, "Delete temp files that are no longer needed.")
|
||||
|
||||
if result.interruptions:
|
||||
state = result.to_state()
|
||||
for interruption in result.interruptions:
|
||||
state.approve(interruption)
|
||||
result = await Runner.run(agent, state)
|
||||
```
|
||||
|
||||
ストリーミング実行では、まず [`stream_events()`][agents.result.RunResultStreaming.stream_events] の消費を完了してから `result.interruptions` を確認し、`result.to_state()` から再開してください。承認フロー全体については、[ヒューマンインザループ](human_in_the_loop.md)を参照してください。
|
||||
|
||||
### サーバー管理の継続
|
||||
|
||||
[`last_response_id`][agents.result.RunResultBase.last_response_id] は、実行から得られた最新のモデルレスポンス ID です。OpenAI Responses API チェーンを継続したい場合は、次のターンで `previous_response_id` として渡してください。
|
||||
|
||||
すでに `to_input_list()`、`session`、または `conversation_id` で会話を継続している場合、通常は `last_response_id` は不要です。複数ステップの実行におけるすべてのモデルレスポンスが必要な場合は、代わりに `raw_responses` を確認してください。
|
||||
|
||||
## ツールとしてのエージェントのメタデータ
|
||||
|
||||
実行結果がネストされた [`Agent.as_tool()`][agents.agent.Agent.as_tool] 実行に由来する場合、[`agent_tool_invocation`][agents.result.RunResultBase.agent_tool_invocation] は外側のツール呼び出しに関する変更不可のメタデータを公開します。
|
||||
|
||||
- `tool_name`
|
||||
- `tool_call_id`
|
||||
- `tool_arguments`
|
||||
|
||||
通常のトップレベル実行では、`agent_tool_invocation` は `None` です。
|
||||
|
||||
これは、ネストされた実行結果を後処理する際に、外側のツール名、呼び出し ID、または生の引数が必要になることがある `custom_output_extractor` 内で特に便利です。関連する `Agent.as_tool()` パターンについては、[ツール](tools.md)を参照してください。
|
||||
|
||||
そのネストされた実行のパース済み構造化入力も必要な場合は、`context_wrapper.tool_input` を読み取ってください。これは、ネストされたツール入力に対して [`RunState`][agents.run_state.RunState] が汎用的にシリアライズするフィールドです。一方、`agent_tool_invocation` は、現在のネストされた呼び出しに対するライブの実行結果アクセサーです。
|
||||
|
||||
## ストリーミングのライフサイクルと診断
|
||||
|
||||
[`RunResultStreaming`][agents.result.RunResultStreaming] は、上記と同じ実行結果サーフェスを継承しますが、ストリーミング固有の制御機能を追加します。
|
||||
|
||||
- セマンティックなストリームイベントを消費するための [`stream_events()`][agents.result.RunResultStreaming.stream_events]
|
||||
- 実行途中でアクティブなエージェントを追跡するための [`current_agent`][agents.result.RunResultStreaming.current_agent]
|
||||
- ストリーミング実行が完全に終了したかどうかを確認するための [`is_complete`][agents.result.RunResultStreaming.is_complete]
|
||||
- 実行を即時または現在のターン後に停止するための [`cancel(...)`][agents.result.RunResultStreaming.cancel]
|
||||
|
||||
非同期イテレーターが終了するまで `stream_events()` を消費し続けてください。そのイテレーターが終了するまで、ストリーミング実行は完了していません。また、`final_output`、`interruptions`、`raw_responses` などの要約プロパティや、セッション永続化の副作用は、目に見える最後のトークンが到着した後もまだ確定中の場合があります。
|
||||
|
||||
`cancel()` を呼び出した場合は、キャンセルとクリーンアップが正しく完了できるように、`stream_events()` を消費し続けてください。
|
||||
|
||||
Python では、ストリーミング用の個別の `completed` プロミスや `error` プロパティは公開されません。終端的なストリーミング失敗は `stream_events()` から例外が送出されることで表面化し、`is_complete` は実行が終端状態に到達したかどうかを反映します。
|
||||
|
||||
### raw レスポンス
|
||||
|
||||
[`raw_responses`][agents.result.RunResultBase.raw_responses] プロパティには、LLM によって生成された [`ModelResponse`][agents.items.ModelResponse] が格納されます。
|
||||
[`raw_responses`][agents.result.RunResultBase.raw_responses] には、実行中に収集された raw モデルレスポンスが含まれます。複数ステップの実行では、ハンドオフをまたいだり、モデル/ツール/モデルのサイクルが繰り返されたりする場合など、複数のレスポンスが生成されることがあります。
|
||||
|
||||
### 元の入力
|
||||
[`last_response_id`][agents.result.RunResultBase.last_response_id] は、`raw_responses` の最後のエントリの ID にすぎません。
|
||||
|
||||
[`input`][agents.result.RunResultBase.input] プロパティには、`run` メソッドに提供した元の入力が格納されます。ほとんどの場合これは必要ありませんが、必要な場合のために利用可能です。
|
||||
### ガードレールの実行結果
|
||||
|
||||
エージェントレベルのガードレールは、[`input_guardrail_results`][agents.result.RunResultBase.input_guardrail_results] および [`output_guardrail_results`][agents.result.RunResultBase.output_guardrail_results] として公開されます。
|
||||
|
||||
ツールガードレールは、[`tool_input_guardrail_results`][agents.result.RunResultBase.tool_input_guardrail_results] および [`tool_output_guardrail_results`][agents.result.RunResultBase.tool_output_guardrail_results] として個別に公開されます。
|
||||
|
||||
これらの配列は実行全体で蓄積されるため、判定のログ記録、追加のガードレールメタデータの保存、または実行がブロックされた理由のデバッグに役立ちます。
|
||||
|
||||
### コンテキストと使用量
|
||||
|
||||
[`context_wrapper`][agents.result.RunResultBase.context_wrapper] は、アプリのコンテキストと、承認、使用量、ネストされた `tool_input` など SDK が管理するランタイムメタデータを公開します。
|
||||
|
||||
使用量は `context_wrapper.usage` で追跡されます。ストリーミング実行では、ストリームの最終チャンクが処理されるまで、使用量の合計値が遅れて反映される場合があります。ラッパーの完全な形状と永続化に関する注意事項については、[コンテキスト管理](context.md)を参照してください。
|
||||
+483
-39
@@ -1,10 +1,14 @@
|
||||
---
|
||||
search:
|
||||
exclude: true
|
||||
---
|
||||
# エージェントの実行
|
||||
|
||||
エージェントは [`Runner`][agents.run.Runner] クラスを使って実行できます。3 つのオプションがあります。
|
||||
[`Runner`][agents.run.Runner] クラスを通じてエージェントを実行できます。3 つの選択肢があります。
|
||||
|
||||
1. [`Runner.run()`][agents.run.Runner.run]:非同期で実行され、[`RunResult`][agents.result.RunResult] を返します。
|
||||
2. [`Runner.run_sync()`][agents.run.Runner.run_sync]:同期メソッドで、内部的には `.run()` を実行します。
|
||||
3. [`Runner.run_streamed()`][agents.run.Runner.run_streamed]:非同期で実行され、[`RunResultStreaming`][agents.result.RunResultStreaming] を返します。LLM をストリーミングモードで呼び出し、受信したイベントをリアルタイムでストリームします。
|
||||
1. [`Runner.run()`][agents.run.Runner.run] は非同期で実行し、[`RunResult`][agents.result.RunResult] を返します。
|
||||
2. [`Runner.run_sync()`][agents.run.Runner.run_sync] は同期メソッドで、内部的には単に `.run()` を実行します。
|
||||
3. [`Runner.run_streamed()`][agents.run.Runner.run_streamed] は非同期で実行し、[`RunResultStreaming`][agents.result.RunResultStreaming] を返します。LLM をストリーミングモードで呼び出し、受信したイベントをそのままストリーミングします。
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner
|
||||
@@ -16,61 +20,267 @@ async def main():
|
||||
print(result.final_output)
|
||||
# Code within the code,
|
||||
# Functions calling themselves,
|
||||
# Infinite loop's dance.
|
||||
# Infinite loop's dance
|
||||
```
|
||||
|
||||
詳細は [results guide](results.md) をご覧ください。
|
||||
詳細は [実行結果ガイド](results.md) を参照してください。
|
||||
|
||||
## エージェントループ
|
||||
## Runner のライフサイクルと設定
|
||||
|
||||
`Runner` の run メソッドを使用する際、開始するエージェントと入力を渡します。入力は文字列(ユーザーメッセージと見なされます)または入力アイテムのリスト(OpenAI Responses API のアイテム)を指定できます。
|
||||
### エージェントループ
|
||||
|
||||
runner は次のようなループを実行します。
|
||||
`Runner` の run メソッドを使用するときは、開始エージェントと入力を渡します。入力には以下を指定できます。
|
||||
|
||||
1. 現在のエージェントと入力で LLM を呼び出します。
|
||||
- 文字列(ユーザーメッセージとして扱われます)
|
||||
- OpenAI Responses API 形式の入力項目のリスト
|
||||
- 中断された実行を再開する場合の [`RunState`][agents.run_state.RunState]
|
||||
|
||||
その後、runner はループを実行します。
|
||||
|
||||
1. 現在のエージェントに対して、現在の入力で LLM を呼び出します。
|
||||
2. LLM が出力を生成します。
|
||||
1. LLM が `final_output` を返した場合、ループは終了し、結果を返します。
|
||||
1. LLM が `final_output` を返した場合、ループは終了し、実行結果を返します。
|
||||
2. LLM がハンドオフを行った場合、現在のエージェントと入力を更新し、ループを再実行します。
|
||||
3. LLM がツール呼び出しを生成した場合、それらのツール呼び出しを実行し、結果を追加してループを再実行します。
|
||||
3. 渡された `max_turns` を超えた場合、[`MaxTurnsExceeded`][agents.exceptions.MaxTurnsExceeded] 例外を発生させます。
|
||||
3. LLM がツール呼び出しを生成した場合、それらのツール呼び出しを実行し、実行結果を追加して、ループを再実行します。
|
||||
3. 渡された `max_turns` を超えた場合、[`MaxTurnsExceeded`][agents.exceptions.MaxTurnsExceeded] 例外を発生させます。このターン制限を無効にするには、`max_turns=None` を渡します。
|
||||
|
||||
!!! note
|
||||
|
||||
LLM の出力が「final output」と見なされるルールは、希望する型のテキスト出力が生成され、ツール呼び出しがない場合です。
|
||||
LLM の出力が「最終出力」と見なされるルールは、目的の型のテキスト出力が生成され、ツール呼び出しがないことです。
|
||||
|
||||
## ストリーミング
|
||||
### ストリーミング
|
||||
|
||||
ストリーミングを利用すると、LLM の実行中にストリーミングイベントを受け取ることができます。ストリームが完了すると、[`RunResultStreaming`][agents.result.RunResultStreaming] に実行に関するすべての情報(新たに生成されたすべての出力を含む)が格納されます。ストリーミングイベントは `.stream_events()` で取得できます。詳細は [streaming guide](streaming.md) をご覧ください。
|
||||
ストリーミングを使用すると、LLM の実行中にストリーミングイベントも受け取れます。ストリームが完了すると、[`RunResultStreaming`][agents.result.RunResultStreaming] には、生成されたすべての新しい出力を含む実行に関する完全な情報が含まれます。ストリーミングイベントには `.stream_events()` を呼び出せます。詳細は [ストリーミングガイド](streaming.md) を参照してください。
|
||||
|
||||
## Run config
|
||||
#### Responses WebSocket トランスポート(任意のヘルパー)
|
||||
|
||||
`run_config` パラメーターでは、エージェント実行のグローバル設定をいくつか構成できます。
|
||||
OpenAI Responses WebSocket トランスポートを有効にしても、通常の `Runner` API を引き続き使用できます。WebSocket セッションヘルパーは接続の再利用に推奨されますが、必須ではありません。
|
||||
|
||||
- [`model`][agents.run.RunConfig.model]:各エージェントの `model` 設定に関わらず、グローバルで使用する LLM モデルを指定できます。
|
||||
- [`model_provider`][agents.run.RunConfig.model_provider]:モデル名を検索するためのモデルプロバイダーで、デフォルトは OpenAI です。
|
||||
- [`model_settings`][agents.run.RunConfig.model_settings]:エージェント固有の設定を上書きします。たとえば、グローバルな `temperature` や `top_p` を設定できます。
|
||||
- [`input_guardrails`][agents.run.RunConfig.input_guardrails], [`output_guardrails`][agents.run.RunConfig.output_guardrails]:すべての実行に含める入力または出力ガードレールのリストです。
|
||||
- [`handoff_input_filter`][agents.run.RunConfig.handoff_input_filter]:すべてのハンドオフに適用するグローバルな入力フィルターです(ハンドオフに既にフィルターがない場合)。入力フィルターを使うと、新しいエージェントに送信する入力を編集できます。詳細は [`Handoff.input_filter`][agents.handoffs.Handoff.input_filter] のドキュメントをご覧ください。
|
||||
- [`tracing_disabled`][agents.run.RunConfig.tracing_disabled]:実行全体の [トレーシング](tracing.md) を無効にできます。
|
||||
- [`trace_include_sensitive_data`][agents.run.RunConfig.trace_include_sensitive_data]:トレースに LLM やツール呼び出しの入出力など、機密性のあるデータを含めるかどうかを設定します。
|
||||
- [`workflow_name`][agents.run.RunConfig.workflow_name], [`trace_id`][agents.run.RunConfig.trace_id], [`group_id`][agents.run.RunConfig.group_id]:実行のトレーシングワークフロー名、トレース ID、トレースグループ ID を設定します。少なくとも `workflow_name` の設定を推奨します。グループ ID は複数の実行にまたがるトレースをリンクするためのオプション項目です。
|
||||
- [`trace_metadata`][agents.run.RunConfig.trace_metadata]:すべてのトレースに含めるメタデータです。
|
||||
これは WebSocket トランスポート上の Responses API であり、[Realtime API](realtime/guide.md) ではありません。
|
||||
|
||||
## 会話・チャットスレッド
|
||||
具体的なモデルオブジェクトやカスタムプロバイダーに関するトランスポート選択ルールと注意点については、[モデル](models/index.md#responses-websocket-transport) を参照してください。
|
||||
|
||||
いずれかの run メソッドを呼び出すと、1 つまたは複数のエージェント(および 1 つまたは複数の LLM 呼び出し)が実行されますが、チャット会話における 1 回の論理的なターンを表します。例:
|
||||
##### パターン 1: セッションヘルパーなし(動作します)
|
||||
|
||||
1. ユーザーのターン:ユーザーがテキストを入力
|
||||
2. Runner の実行:最初のエージェントが LLM を呼び出し、ツールを実行し、2 番目のエージェントにハンドオフ、2 番目のエージェントがさらにツールを実行し、出力を生成
|
||||
WebSocket トランスポートだけを使用したく、SDK に共有プロバイダー / セッションを管理させる必要がない場合に使用します。
|
||||
|
||||
エージェントの実行が終わったら、ユーザーに何を表示するか選択できます。たとえば、エージェントが生成したすべての新しいアイテムをユーザーに見せることも、最終出力だけを見せることもできます。いずれの場合も、ユーザーが追加の質問をした場合は、再度 run メソッドを呼び出せます。
|
||||
```python
|
||||
import asyncio
|
||||
|
||||
次のターンの入力を取得するには、ベースの [`RunResultBase.to_input_list()`][agents.result.RunResultBase.to_input_list] メソッドを使用できます。
|
||||
from agents import Agent, Runner, set_default_openai_responses_transport
|
||||
|
||||
|
||||
async def main():
|
||||
set_default_openai_responses_transport("websocket")
|
||||
|
||||
agent = Agent(name="Assistant", instructions="Be concise.")
|
||||
result = Runner.run_streamed(agent, "Summarize recursion in one sentence.")
|
||||
|
||||
async for event in result.stream_events():
|
||||
if event.type == "raw_response_event":
|
||||
continue
|
||||
print(event.type)
|
||||
|
||||
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
このパターンは単発の実行には問題ありません。`Runner.run()` / `Runner.run_streamed()` を繰り返し呼び出す場合、同じ `RunConfig` / プロバイダーインスタンスを手動で再利用しない限り、各実行で再接続される可能性があります。
|
||||
|
||||
##### パターン 2: `responses_websocket_session()` の使用(複数ターンの再利用に推奨)
|
||||
|
||||
複数の実行にわたって共有の WebSocket 対応プロバイダーと `RunConfig` を使用したい場合は、[`responses_websocket_session()`][agents.responses_websocket_session] を使用します(同じ `run_config` を継承するネストされた agent-as-tool 呼び出しを含みます)。
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
|
||||
from agents import Agent, responses_websocket_session
|
||||
|
||||
|
||||
async def main():
|
||||
agent = Agent(name="Assistant", instructions="Be concise.")
|
||||
|
||||
async with responses_websocket_session(
|
||||
responses_websocket_options={"ping_interval": 20.0, "ping_timeout": 60.0},
|
||||
) as ws:
|
||||
first = ws.run_streamed(agent, "Say hello in one short sentence.")
|
||||
async for _event in first.stream_events():
|
||||
pass
|
||||
|
||||
second = ws.run_streamed(
|
||||
agent,
|
||||
"Now say goodbye.",
|
||||
previous_response_id=first.last_response_id,
|
||||
)
|
||||
async for _event in second.stream_events():
|
||||
pass
|
||||
|
||||
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
コンテキストを抜ける前に、ストリーミングされた実行結果の消費を完了してください。WebSocket リクエストがまだ実行中の状態でコンテキストを抜けると、共有接続が強制的に閉じられる可能性があります。
|
||||
|
||||
長い推論ターンで WebSocket の keepalive タイムアウトに達する場合は、`ping_timeout` を増やすか、`ping_timeout=None` を設定してハートビートタイムアウトを無効にしてください。WebSocket のレイテンシより信頼性が重要な実行では、HTTP/SSE トランスポートを使用してください。
|
||||
|
||||
### 実行設定
|
||||
|
||||
`run_config` パラメーターを使用すると、エージェント実行に関するいくつかのグローバル設定を構成できます。
|
||||
|
||||
#### 一般的な実行設定カテゴリー
|
||||
|
||||
各エージェント定義を変更せずに単一の実行の動作を上書きするには、`RunConfig` を使用します。
|
||||
|
||||
##### モデル、プロバイダー、セッションのデフォルト
|
||||
|
||||
- [`model`][agents.run.RunConfig.model]: 各 Agent が持つ `model` に関係なく、使用するグローバルな LLM モデルを設定できます。
|
||||
- [`model_provider`][agents.run.RunConfig.model_provider]: モデル名の検索に使用するモデルプロバイダーで、デフォルトは OpenAI です。
|
||||
- [`model_settings`][agents.run.RunConfig.model_settings]: エージェント固有の設定を上書きします。たとえば、グローバルな `temperature` や `top_p` を設定できます。
|
||||
- [`session_settings`][agents.run.RunConfig.session_settings]: 実行中に履歴を取得する際のセッションレベルのデフォルト(たとえば `SessionSettings(limit=...)`)を上書きします。
|
||||
- [`session_input_callback`][agents.run.RunConfig.session_input_callback]: Sessions を使用する際、各ターンの前に新しいユーザー入力をセッション履歴とどのようにマージするかをカスタマイズします。コールバックは同期でも非同期でもかまいません。
|
||||
|
||||
##### ガードレール、ハンドオフ、モデル入力の整形
|
||||
|
||||
- [`input_guardrails`][agents.run.RunConfig.input_guardrails], [`output_guardrails`][agents.run.RunConfig.output_guardrails]: すべての実行に含める入力または出力ガードレールのリストです。
|
||||
- [`handoff_input_filter`][agents.run.RunConfig.handoff_input_filter]: ハンドオフにまだ入力フィルターがない場合に、すべてのハンドオフに適用するグローバル入力フィルターです。入力フィルターを使用すると、新しいエージェントに送信される入力を編集できます。詳細は [`Handoff.input_filter`][agents.handoffs.Handoff.input_filter] のドキュメントを参照してください。
|
||||
- [`nest_handoff_history`][agents.run.RunConfig.nest_handoff_history]: 次のエージェントを呼び出す前に、以前のトランスクリプトを単一のアシスタントメッセージに折りたたむオプトインのベータ機能です。ネストされたハンドオフを安定化させている間、この機能はデフォルトで無効です。有効にするには `True` に設定し、未加工のトランスクリプトをそのまま渡すには `False` のままにします。すべての [Runner メソッド][agents.run.Runner] は、渡されない場合に自動的に `RunConfig` を作成するため、クイックスタートとコード例ではデフォルトでオフのままになり、明示的な [`Handoff.input_filter`][agents.handoffs.Handoff.input_filter] コールバックは引き続きこれを上書きします。個別のハンドオフは [`Handoff.nest_handoff_history`][agents.handoffs.Handoff.nest_handoff_history] を通じてこの設定を上書きできます。
|
||||
- [`handoff_history_mapper`][agents.run.RunConfig.handoff_history_mapper]: `nest_handoff_history` にオプトインした場合に、正規化されたトランスクリプト(履歴 + ハンドオフ項目)を受け取る任意の callable です。次のエージェントに転送する入力項目の正確なリストを返す必要があり、完全なハンドオフフィルターを書かずに組み込みの要約を置き換えられます。
|
||||
- [`call_model_input_filter`][agents.run.RunConfig.call_model_input_filter]: モデル呼び出しの直前に、完全に準備されたモデル入力(instructions と入力項目)を編集するフックです。たとえば、履歴をトリミングしたり、システムプロンプトを挿入したりできます。
|
||||
- [`reasoning_item_id_policy`][agents.run.RunConfig.reasoning_item_id_policy]: runner が以前の出力を次ターンのモデル入力に変換するときに、推論項目 ID を保持するか省略するかを制御します。
|
||||
|
||||
##### トレーシングと可観測性
|
||||
|
||||
- [`tracing_disabled`][agents.run.RunConfig.tracing_disabled]: 実行全体で [トレーシング](tracing.md) を無効にできます。
|
||||
- [`tracing`][agents.run.RunConfig.tracing]: 実行ごとのトレーシング API キーなど、トレースエクスポート設定を上書きするために [`TracingConfig`][agents.tracing.TracingConfig] を渡します。
|
||||
- [`trace_include_sensitive_data`][agents.run.RunConfig.trace_include_sensitive_data]: トレースに、LLM やツール呼び出しの入力 / 出力など、潜在的に機密性の高いデータを含めるかどうかを設定します。
|
||||
- [`workflow_name`][agents.run.RunConfig.workflow_name], [`trace_id`][agents.run.RunConfig.trace_id], [`group_id`][agents.run.RunConfig.group_id]: 実行のトレーシングワークフロー名、トレース ID、トレースグループ ID を設定します。少なくとも `workflow_name` を設定することをおすすめします。グループ ID は、複数の実行間でトレースを関連付けるための任意フィールドです。
|
||||
- [`trace_metadata`][agents.run.RunConfig.trace_metadata]: すべてのトレースに含めるメタデータです。
|
||||
|
||||
##### ツール実行、承認、ツールエラーの動作
|
||||
|
||||
- [`tool_execution`][agents.run.RunConfig.tool_execution]: 一度に実行する関数ツール数を制限するなど、ローカルツール呼び出しに対する SDK 側の実行動作を設定します。
|
||||
- [`tool_error_formatter`][agents.run.RunConfig.tool_error_formatter]: 承認フロー中にツール呼び出しが拒否された場合に、モデルに表示されるメッセージをカスタマイズします。
|
||||
|
||||
ネストされたハンドオフは、オプトインのベータ機能として利用できます。`RunConfig(nest_handoff_history=True)` を渡すか、特定のハンドオフで `handoff(..., nest_handoff_history=True)` を設定すると、折りたたみトランスクリプト動作を有効にできます。未加工のトランスクリプトを保持したい場合(デフォルト)は、フラグを未設定のままにするか、必要に応じて会話を正確に転送する `handoff_input_filter`(または `handoff_history_mapper`)を指定します。カスタムマッパーを書かずに、生成される要約で使用されるラッパーテキストを変更するには、[`set_conversation_history_wrappers`][agents.handoffs.set_conversation_history_wrappers] を呼び出します(デフォルトに戻すには [`reset_conversation_history_wrappers`][agents.handoffs.reset_conversation_history_wrappers])。
|
||||
|
||||
#### 実行設定の詳細
|
||||
|
||||
##### `tool_execution`
|
||||
|
||||
実行に対してローカル関数ツールの同時実行数を SDK に制限させたい場合は、`tool_execution` を使用します。
|
||||
|
||||
```python
|
||||
from agents import Agent, RunConfig, Runner, ToolExecutionConfig
|
||||
|
||||
agent = Agent(name="Assistant", tools=[...])
|
||||
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
"Run the required tool calls.",
|
||||
run_config=RunConfig(
|
||||
tool_execution=ToolExecutionConfig(max_function_tool_concurrency=2),
|
||||
),
|
||||
)
|
||||
```
|
||||
|
||||
`max_function_tool_concurrency=None` はデフォルトの動作を維持します。モデルが 1 ターンで複数の関数ツール呼び出しを生成した場合、SDK は生成されたすべてのローカル関数ツール呼び出しを開始します。整数値を設定すると、それらのローカル関数ツールのうち同時に実行される数に上限を設けられます。
|
||||
|
||||
これはプロバイダー側の [`ModelSettings.parallel_tool_calls`][agents.model_settings.ModelSettings.parallel_tool_calls] とは別のものです。`parallel_tool_calls` は、モデルが 1 つのレスポンスで複数のツール呼び出しを生成できるかどうかを制御します。`tool_execution.max_function_tool_concurrency` は、モデルがそれらを生成した後に、SDK がローカル関数ツール呼び出しをどのように実行するかを制御します。
|
||||
|
||||
##### `tool_error_formatter`
|
||||
|
||||
承認フローでツール呼び出しが拒否されたときにモデルへ返されるメッセージをカスタマイズするには、`tool_error_formatter` を使用します。
|
||||
|
||||
フォーマッターは、以下を含む [`ToolErrorFormatterArgs`][agents.run_config.ToolErrorFormatterArgs] を受け取ります。
|
||||
|
||||
- `kind`: エラーカテゴリーです。現在は `"approval_rejected"` です。
|
||||
- `tool_type`: ツールランタイム(`"function"`、`"computer"`、`"shell"`、`"apply_patch"`、または `"custom"`)です。
|
||||
- `tool_name`: ツール名です。
|
||||
- `call_id`: ツール呼び出し ID です。
|
||||
- `default_message`: SDK のデフォルトのモデル表示メッセージです。
|
||||
- `run_context`: アクティブな実行コンテキストラッパーです。
|
||||
|
||||
メッセージを置き換えるには文字列を返し、SDK のデフォルトを使用するには `None` を返します。
|
||||
|
||||
```python
|
||||
from agents import Agent, RunConfig, Runner, ToolErrorFormatterArgs
|
||||
|
||||
|
||||
def format_rejection(args: ToolErrorFormatterArgs[None]) -> str | None:
|
||||
if args.kind == "approval_rejected":
|
||||
return (
|
||||
f"Tool call '{args.tool_name}' was rejected by a human reviewer. "
|
||||
"Ask for confirmation or propose a safer alternative."
|
||||
)
|
||||
return None
|
||||
|
||||
|
||||
agent = Agent(name="Assistant")
|
||||
result = Runner.run_sync(
|
||||
agent,
|
||||
"Please delete the production database.",
|
||||
run_config=RunConfig(tool_error_formatter=format_rejection),
|
||||
)
|
||||
```
|
||||
|
||||
##### `reasoning_item_id_policy`
|
||||
|
||||
`reasoning_item_id_policy` は、runner が履歴を引き継ぐとき(たとえば `RunResult.to_input_list()` やセッションに基づく実行を使用する場合)に、推論項目を次ターンのモデル入力へどのように変換するかを制御します。
|
||||
|
||||
- `None` または `"preserve"`(デフォルト): 推論項目 ID を保持します。
|
||||
- `"omit"`: 生成される次ターンの入力から推論項目 ID を取り除きます。
|
||||
|
||||
`"omit"` は主に、推論項目が `id` 付きで送信されているものの、必須の後続項目がない場合に発生する Responses API 400 エラーの一群に対する、オプトインの緩和策として使用します(例: `Item 'rs_...' of type 'reasoning' was provided without its required following item.`)。
|
||||
|
||||
これは、SDK が以前の出力からフォローアップ入力を構築するマルチターンのエージェント実行で発生する可能性があります(セッション永続化、サーバー管理の会話差分、ストリーミング / 非ストリーミングのフォローアップターン、再開パスを含みます)。推論項目 ID が保持されている一方で、プロバイダーがその ID を対応する後続項目とペアのままにすることを要求する場合です。
|
||||
|
||||
`reasoning_item_id_policy="omit"` を設定すると、推論コンテンツは保持したまま推論項目の `id` を取り除くため、SDK が生成するフォローアップ入力でその API 不変条件がトリガーされることを回避できます。
|
||||
|
||||
スコープに関する注記:
|
||||
|
||||
- これは、SDK がフォローアップ入力を構築するときに SDK によって生成 / 転送される推論項目のみを変更します。
|
||||
- ユーザーが指定した初期入力項目は書き換えません。
|
||||
- `call_model_input_filter` は、このポリシー適用後でも意図的に推論 ID を再導入できます。
|
||||
|
||||
## 状態と会話管理
|
||||
|
||||
### メモリ戦略の選択
|
||||
|
||||
状態を次のターンへ引き継ぐ一般的な方法は 4 つあります。
|
||||
|
||||
| 戦略 | 状態の保存場所 | 最適な用途 | 次のターンで渡すもの |
|
||||
| --- | --- | --- | --- |
|
||||
| `result.to_input_list()` | アプリのメモリ | 小規模なチャットループ、完全な手動制御、任意のプロバイダー | `result.to_input_list()` からのリストに次のユーザーメッセージを加えたもの |
|
||||
| `session` | ストレージと SDK | 永続的なチャット状態、再開可能な実行、カスタムストア | 同じ `session` インスタンス、または同じストアを指す別のインスタンス |
|
||||
| `conversation_id` | OpenAI Conversations API | ワーカーやサービス間で共有したい名前付きのサーバー側会話 | 同じ `conversation_id` と新しいユーザーターンのみ |
|
||||
| `previous_response_id` | OpenAI Responses API | 会話リソースを作成しない、軽量なサーバー管理の継続 | `result.last_response_id` と新しいユーザーターンのみ |
|
||||
|
||||
`result.to_input_list()` と `session` はクライアント管理です。`conversation_id` と `previous_response_id` は OpenAI 管理で、OpenAI Responses API を使用している場合にのみ適用されます。ほとんどのアプリケーションでは、会話ごとに 1 つの永続化戦略を選択します。意図的に両方のレイヤーを調整している場合を除き、クライアント管理の履歴と OpenAI 管理の状態を混在させると、コンテキストが重複する可能性があります。
|
||||
|
||||
!!! note
|
||||
|
||||
セッション永続化は、同じ実行内でサーバー管理の会話設定
|
||||
(`conversation_id`、`previous_response_id`、または `auto_previous_response_id`)と
|
||||
組み合わせることはできません。呼び出しごとに 1 つのアプローチを選択してください。
|
||||
|
||||
### 会話 / チャットスレッド
|
||||
|
||||
いずれかの run メソッドを呼び出すと、1 つ以上のエージェントが実行される(したがって 1 回以上の LLM 呼び出しが行われる)場合がありますが、チャット会話における単一の論理ターンを表します。例:
|
||||
|
||||
1. ユーザーターン: ユーザーがテキストを入力します
|
||||
2. Runner の実行: 最初のエージェントが LLM を呼び出し、ツールを実行し、2 番目のエージェントへハンドオフし、2 番目のエージェントがさらにツールを実行してから出力を生成します。
|
||||
|
||||
エージェント実行の終了時に、ユーザーに何を表示するかを選択できます。たとえば、エージェントによって生成されたすべての新しい項目をユーザーに表示することも、最終出力だけを表示することもできます。いずれの場合でも、その後ユーザーが追加の質問をする可能性があり、その場合は run メソッドを再度呼び出せます。
|
||||
|
||||
#### 手動の会話管理
|
||||
|
||||
[`RunResultBase.to_input_list()`][agents.result.RunResultBase.to_input_list] メソッドを使用して次のターンの入力を取得し、会話履歴を手動で管理できます。
|
||||
|
||||
```python
|
||||
async def main():
|
||||
agent = Agent(name="Assistant", instructions="Reply very concisely.")
|
||||
|
||||
thread_id = "thread_123" # Example thread ID
|
||||
with trace(workflow_name="Conversation", group_id=thread_id):
|
||||
# First turn
|
||||
result = await Runner.run(agent, "What city is the Golden Gate Bridge in?")
|
||||
@@ -84,12 +294,246 @@ async def main():
|
||||
# California
|
||||
```
|
||||
|
||||
#### セッションによる自動会話管理
|
||||
|
||||
より簡単な方法として、[Sessions](sessions/index.md) を使用すると、`.to_input_list()` を手動で呼び出すことなく会話履歴を自動的に処理できます。
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner, SQLiteSession
|
||||
|
||||
async def main():
|
||||
agent = Agent(name="Assistant", instructions="Reply very concisely.")
|
||||
|
||||
# Create session instance
|
||||
session = SQLiteSession("conversation_123")
|
||||
|
||||
thread_id = "thread_123" # Example thread ID
|
||||
with trace(workflow_name="Conversation", group_id=thread_id):
|
||||
# First turn
|
||||
result = await Runner.run(agent, "What city is the Golden Gate Bridge in?", session=session)
|
||||
print(result.final_output)
|
||||
# San Francisco
|
||||
|
||||
# Second turn - agent automatically remembers previous context
|
||||
result = await Runner.run(agent, "What state is it in?", session=session)
|
||||
print(result.final_output)
|
||||
# California
|
||||
```
|
||||
|
||||
Sessions は自動的に以下を行います。
|
||||
|
||||
- 各実行の前に会話履歴を取得します
|
||||
- 各実行の後に新しいメッセージを保存します
|
||||
- 異なるセッション ID ごとに別々の会話を維持します
|
||||
|
||||
詳細は [Sessions ドキュメント](sessions/index.md) を参照してください。
|
||||
|
||||
|
||||
#### サーバー管理の会話
|
||||
|
||||
`to_input_list()` や `Sessions` でローカルに処理する代わりに、OpenAI の会話状態機能にサーバー側で会話状態を管理させることもできます。これにより、過去のすべてのメッセージを手動で再送信することなく、会話履歴を保持できます。以下のどちらのサーバー管理アプローチでも、各リクエストでは新しいターンの入力だけを渡し、保存した ID を再利用します。詳細は [OpenAI Conversation state ガイド](https://platform.openai.com/docs/guides/conversation-state?api-mode=responses) を参照してください。
|
||||
|
||||
OpenAI は、ターン間で状態を追跡する 2 つの方法を提供しています。
|
||||
|
||||
##### 1. `conversation_id` の使用
|
||||
|
||||
まず OpenAI Conversations API を使用して会話を作成し、その後のすべての呼び出しでその ID を再利用します。
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner
|
||||
from openai import AsyncOpenAI
|
||||
|
||||
client = AsyncOpenAI()
|
||||
|
||||
async def main():
|
||||
agent = Agent(name="Assistant", instructions="Reply very concisely.")
|
||||
|
||||
# Create a server-managed conversation
|
||||
conversation = await client.conversations.create()
|
||||
conv_id = conversation.id
|
||||
|
||||
while True:
|
||||
user_input = input("You: ")
|
||||
result = await Runner.run(agent, user_input, conversation_id=conv_id)
|
||||
print(f"Assistant: {result.final_output}")
|
||||
```
|
||||
|
||||
##### 2. `previous_response_id` の使用
|
||||
|
||||
もう 1 つの選択肢は **レスポンスチェーン** で、各ターンを前のターンのレスポンス ID に明示的にリンクします。
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner
|
||||
|
||||
async def main():
|
||||
agent = Agent(name="Assistant", instructions="Reply very concisely.")
|
||||
|
||||
previous_response_id = None
|
||||
|
||||
while True:
|
||||
user_input = input("You: ")
|
||||
|
||||
# Setting auto_previous_response_id=True enables response chaining automatically
|
||||
# for the first turn, even when there's no actual previous response ID yet.
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
user_input,
|
||||
previous_response_id=previous_response_id,
|
||||
auto_previous_response_id=True,
|
||||
)
|
||||
previous_response_id = result.last_response_id
|
||||
print(f"Assistant: {result.final_output}")
|
||||
```
|
||||
|
||||
実行が承認のために一時停止し、[`RunState`][agents.run_state.RunState] から再開した場合、
|
||||
SDK は保存された `conversation_id` / `previous_response_id` / `auto_previous_response_id`
|
||||
設定を保持するため、再開されたターンは同じサーバー管理の会話で続行されます。
|
||||
|
||||
`conversation_id` と `previous_response_id` は同時に使用できません。システム間で共有できる名前付きの会話リソースが必要な場合は `conversation_id` を使用します。あるターンから次のターンへ最も軽量な Responses API の継続基本コンポーネントが必要な場合は `previous_response_id` を使用します。
|
||||
|
||||
!!! note
|
||||
|
||||
SDK は `conversation_locked` エラーをバックオフ付きで自動的に再試行します。サーバー管理の
|
||||
会話実行では、再試行前に内部の会話トラッカー入力を巻き戻し、同じ準備済み項目を
|
||||
クリーンに再送信できるようにします。
|
||||
|
||||
ローカルセッションベースの実行(`conversation_id`、
|
||||
`previous_response_id`、または `auto_previous_response_id` と組み合わせることはできません)では、SDK は再試行後の重複した履歴エントリを減らすために、
|
||||
直近に永続化された入力項目のベストエフォートなロールバックも実行します。
|
||||
|
||||
この互換性再試行は、`ModelSettings.retry` を設定していない場合でも発生します。モデルリクエストに対するより広範なオプトインの再試行動作については、[Runner 管理の再試行](models/index.md#runner-managed-retries) を参照してください。
|
||||
|
||||
## フックとカスタマイズ
|
||||
|
||||
### モデル呼び出し入力フィルター
|
||||
|
||||
モデル呼び出しの直前にモデル入力を編集するには、`call_model_input_filter` を使用します。このフックは現在のエージェント、コンテキスト、結合された入力項目(存在する場合はセッション履歴を含む)を受け取り、新しい `ModelInputData` を返します。
|
||||
|
||||
戻り値は [`ModelInputData`][agents.run.ModelInputData] オブジェクトである必要があります。その `input` フィールドは必須で、入力項目のリストでなければなりません。それ以外の形状を返すと `UserError` が発生します。
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner, RunConfig
|
||||
from agents.run import CallModelData, ModelInputData
|
||||
|
||||
def drop_old_messages(data: CallModelData[None]) -> ModelInputData:
|
||||
# Keep only the last 5 items and preserve existing instructions.
|
||||
trimmed = data.model_data.input[-5:]
|
||||
return ModelInputData(input=trimmed, instructions=data.model_data.instructions)
|
||||
|
||||
agent = Agent(name="Assistant", instructions="Answer concisely.")
|
||||
result = Runner.run_sync(
|
||||
agent,
|
||||
"Explain quines",
|
||||
run_config=RunConfig(call_model_input_filter=drop_old_messages),
|
||||
)
|
||||
```
|
||||
|
||||
runner は準備済み入力リストのコピーをフックに渡すため、呼び出し元の元のリストをその場で変更することなく、トリミング、置換、並べ替えができます。
|
||||
|
||||
セッションを使用している場合、`call_model_input_filter` はセッション履歴がすでに読み込まれ、現在のターンとマージされた後に実行されます。その前段のマージ手順自体をカスタマイズしたい場合は、[`session_input_callback`][agents.run.RunConfig.session_input_callback] を使用してください。
|
||||
|
||||
`conversation_id`、`previous_response_id`、または `auto_previous_response_id` を使用して OpenAI のサーバー管理の会話状態を使用している場合、このフックは次の Responses API 呼び出し用に準備されたペイロード上で実行されます。そのペイロードは、以前の履歴全体の再生ではなく、すでに新しいターンの差分のみを表している場合があります。返した項目だけが、そのサーバー管理の継続に対して送信済みとしてマークされます。
|
||||
|
||||
機密データの編集、長い履歴のトリミング、追加のシステムガイダンスの挿入を行うには、実行ごとに `run_config` でこのフックを設定します。
|
||||
|
||||
## エラーと復旧
|
||||
|
||||
### エラーハンドラー
|
||||
|
||||
すべての `Runner` エントリーポイントは、エラー種別をキーとする dict である `error_handlers` を受け取ります。サポートされるキーは `"max_turns"` と `"model_refusal"` です。`MaxTurnsExceeded` または `ModelRefusalError` を発生させる代わりに、制御された最終出力を返したい場合に使用します。
|
||||
|
||||
```python
|
||||
from agents import (
|
||||
Agent,
|
||||
RunErrorHandlerInput,
|
||||
RunErrorHandlerResult,
|
||||
Runner,
|
||||
)
|
||||
|
||||
agent = Agent(name="Assistant", instructions="Be concise.")
|
||||
|
||||
|
||||
def on_max_turns(_data: RunErrorHandlerInput[None]) -> RunErrorHandlerResult:
|
||||
return RunErrorHandlerResult(
|
||||
final_output="I couldn't finish within the turn limit. Please narrow the request.",
|
||||
include_in_history=False,
|
||||
)
|
||||
|
||||
|
||||
result = Runner.run_sync(
|
||||
agent,
|
||||
"Analyze this long transcript",
|
||||
max_turns=3,
|
||||
error_handlers={"max_turns": on_max_turns},
|
||||
)
|
||||
print(result.final_output)
|
||||
```
|
||||
|
||||
フォールバック出力を会話履歴に追加したくない場合は、`include_in_history=False` を設定します。
|
||||
|
||||
モデル拒否が `ModelRefusalError` で実行を終了するのではなく、アプリケーション固有のフォールバックを生成すべき場合は、`"model_refusal"` を使用します。
|
||||
|
||||
```python
|
||||
from pydantic import BaseModel
|
||||
|
||||
from agents import Agent, ModelRefusalError, RunErrorHandlerInput, Runner
|
||||
|
||||
|
||||
class Recipe(BaseModel):
|
||||
ingredients: list[str]
|
||||
refusal_reason: str | None = None
|
||||
|
||||
|
||||
def on_model_refusal(data: RunErrorHandlerInput[None]) -> Recipe:
|
||||
assert isinstance(data.error, ModelRefusalError)
|
||||
return Recipe(ingredients=[], refusal_reason=data.error.refusal)
|
||||
|
||||
|
||||
agent = Agent(
|
||||
name="Recipe assistant",
|
||||
instructions="Return a structured recipe.",
|
||||
output_type=Recipe,
|
||||
)
|
||||
|
||||
result = Runner.run_sync(
|
||||
agent,
|
||||
"Make me something unsafe.",
|
||||
error_handlers={"model_refusal": on_model_refusal},
|
||||
)
|
||||
print(result.final_output)
|
||||
```
|
||||
|
||||
## 耐久実行インテグレーションと human-in-the-loop
|
||||
|
||||
ツール承認の一時停止 / 再開パターンについては、専用の [Human-in-the-loop ガイド](human_in_the_loop.md) から始めてください。
|
||||
以下のインテグレーションは、実行が長い待機、再試行、またはプロセス再起動をまたぐ可能性がある場合の耐久性のあるオーケストレーション向けです。
|
||||
|
||||
### Dapr
|
||||
|
||||
Agents SDK の [Dapr](https://dapr.io) Diagrid インテグレーションを使用すると、human-in-the-loop サポート付きで障害から自動的に復旧する、耐久性のある長時間実行エージェントを実行できます。Dapr はベンダーニュートラルな [CNCF](https://cncf.io) ワークフローオーケストレーターです。Dapr と OpenAI エージェントの始め方は [こちら](https://docs.diagrid.io/getting-started/quickstarts/ai-agents/?agentframework=openai) です。
|
||||
|
||||
### Temporal
|
||||
|
||||
Agents SDK の [Temporal](https://temporal.io/) インテグレーションを使用すると、human-in-the-loop タスクを含む、耐久性のある長時間実行ワークフローを実行できます。Temporal と Agents SDK が連携して長時間実行タスクを完了するデモは [この動画](https://www.youtube.com/watch?v=fFBZqzT4DD8) で確認でき、[ドキュメントはこちら](https://github.com/temporalio/sdk-python/tree/main/temporalio/contrib/openai_agents) です。
|
||||
|
||||
### Restate
|
||||
|
||||
Agents SDK の [Restate](https://restate.dev/) インテグレーションを使用すると、人間による承認、ハンドオフ、セッション管理を含む、軽量で耐久性のあるエージェントを利用できます。このインテグレーションは依存関係として Restate の単一バイナリランタイムを必要とし、エージェントをプロセス / コンテナまたはサーバーレス関数として実行することをサポートします。
|
||||
詳細は [概要](https://www.restate.dev/blog/durable-orchestration-for-ai-agents-with-restate-and-openai-sdk) を読むか、[ドキュメント](https://docs.restate.dev/ai) を参照してください。
|
||||
|
||||
### DBOS
|
||||
|
||||
Agents SDK の [DBOS](https://dbos.dev/) インテグレーションを使用すると、障害や再起動をまたいで進捗を保持する信頼性の高いエージェントを実行できます。長時間実行エージェント、human-in-the-loop ワークフロー、ハンドオフをサポートします。同期メソッドと非同期メソッドの両方をサポートします。このインテグレーションに必要なのは SQLite または Postgres データベースのみです。詳細はインテグレーションの [repo](https://github.com/dbos-inc/dbos-openai-agents) と [ドキュメント](https://docs.dbos.dev/integrations/openai-agents) を参照してください。
|
||||
|
||||
## 例外
|
||||
|
||||
SDK は特定のケースで例外を発生させます。全リストは [`agents.exceptions`][] にあります。概要は以下の通りです。
|
||||
SDK は特定のケースで例外を発生させます。完全な一覧は [`agents.exceptions`][] にあります。概要は以下のとおりです。
|
||||
|
||||
- [`AgentsException`][agents.exceptions.AgentsException]:SDK で発生するすべての例外の基底クラスです。
|
||||
- [`MaxTurnsExceeded`][agents.exceptions.MaxTurnsExceeded]:run メソッドに渡した `max_turns` を超えた場合に発生します。
|
||||
- [`ModelBehaviorError`][agents.exceptions.ModelBehaviorError]:モデルが不正な出力(例:不正な JSON や存在しないツールの使用)を生成した場合に発生します。
|
||||
- [`UserError`][agents.exceptions.UserError]:SDK を使用する際に、あなた(SDK を使ってコードを書く人)がエラーを起こした場合に発生します。
|
||||
- [`InputGuardrailTripwireTriggered`][agents.exceptions.InputGuardrailTripwireTriggered], [`OutputGuardrailTripwireTriggered`][agents.exceptions.OutputGuardrailTripwireTriggered]:[ガードレール](guardrails.md) が作動した場合に発生します。
|
||||
- [`AgentsException`][agents.exceptions.AgentsException]: SDK 内で発生するすべての例外の基底クラスです。他のすべての具体的な例外の派生元となる汎用型として機能します。
|
||||
- [`MaxTurnsExceeded`][agents.exceptions.MaxTurnsExceeded]: この例外は、エージェントの実行が `Runner.run`、`Runner.run_sync`、または `Runner.run_streamed` メソッドに渡された `max_turns` 制限を超えた場合に発生します。エージェントが指定された相互作用ターン数内にタスクを完了できなかったことを示します。制限を無効にするには `max_turns=None` を設定します。
|
||||
- [`ModelBehaviorError`][agents.exceptions.ModelBehaviorError]: この例外は、基盤となるモデル(LLM)が予期しない、または無効な出力を生成した場合に発生します。これには以下が含まれる場合があります。
|
||||
- 不正な形式の JSON: モデルがツール呼び出しまたは直接出力で不正な形式の JSON 構造を提供した場合、特に特定の `output_type` が定義されている場合です。
|
||||
- 予期しないツール関連の失敗: モデルが期待どおりにツールを使用できなかった場合です
|
||||
- [`ToolTimeoutError`][agents.exceptions.ToolTimeoutError]: この例外は、関数ツール呼び出しが設定されたタイムアウトを超え、そのツールが `timeout_behavior="raise_exception"` を使用している場合に発生します。
|
||||
- [`UserError`][agents.exceptions.UserError]: この例外は、あなた(SDK を使用してコードを書いている人)が SDK の使用中にエラーを起こした場合に発生します。通常、誤ったコード実装、無効な設定、または SDK の API の誤用が原因です。
|
||||
- [`InputGuardrailTripwireTriggered`][agents.exceptions.InputGuardrailTripwireTriggered], [`OutputGuardrailTripwireTriggered`][agents.exceptions.OutputGuardrailTripwireTriggered]: この例外は、それぞれ入力ガードレールまたは出力ガードレールの条件が満たされた場合に発生します。入力ガードレールは処理前に受信メッセージをチェックし、出力ガードレールは配信前にエージェントの最終レスポンスをチェックします。
|
||||
@@ -0,0 +1,141 @@
|
||||
---
|
||||
search:
|
||||
exclude: true
|
||||
---
|
||||
# サンドボックスクライアント
|
||||
|
||||
このページでは、サンドボックスでの作業をどこで実行するかを選択します。ほとんどの場合、`SandboxAgent` の定義は同じままにし、サンドボックスクライアントとクライアント固有のオプションを [`SandboxRunConfig`][agents.run_config.SandboxRunConfig] で変更します。
|
||||
|
||||
!!! warning "ベータ機能"
|
||||
|
||||
サンドボックスエージェントはベータ版です。一般提供までに API の詳細、デフォルト値、サポートされる機能が変更される可能性があります。また、時間とともにより高度な機能が追加される見込みです。
|
||||
|
||||
## 判断ガイド
|
||||
|
||||
<div class="sandbox-nowrap-first-column-table" markdown="1">
|
||||
|
||||
| 目的 | まず使うもの | 理由 |
|
||||
| --- | --- | --- |
|
||||
| macOS または Linux での最速のローカル反復 | `UnixLocalSandboxClient` | 追加インストール不要で、シンプルなローカルファイルシステム開発ができます。 |
|
||||
| 基本的なコンテナ分離 | `DockerSandboxClient` | 特定のイメージを使って Docker 内で作業を実行します。 |
|
||||
| ホスト型実行または本番環境スタイルの分離 | ホスト型サンドボックスクライアント | ワークスペース境界をプロバイダー管理環境へ移します。 |
|
||||
|
||||
</div>
|
||||
|
||||
## ローカルクライアント
|
||||
|
||||
ほとんどのユーザーは、これら 2 つのサンドボックスクライアントのいずれかから始めることをおすすめします。
|
||||
|
||||
<div class="sandbox-nowrap-first-column-table" markdown="1">
|
||||
|
||||
| クライアント | インストール | 選ぶ場面 | 例 |
|
||||
| --- | --- | --- | --- |
|
||||
| `UnixLocalSandboxClient` | なし | macOS または Linux で最速のローカル反復が必要な場合。ローカル開発の既定として適しています。 | [Unix-local スターター](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/unix_local_runner.py) |
|
||||
| `DockerSandboxClient` | `openai-agents[docker]` | コンテナ分離、またはローカルで同等性を保つための特定のイメージが必要な場合。 | [Docker スターター](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/docker/docker_runner.py) |
|
||||
|
||||
</div>
|
||||
|
||||
Unix-local は、ローカルファイルシステムに対して開発を始める最も簡単な方法です。より強い環境分離や本番環境スタイルの同等性が必要になったら、Docker またはホスト型プロバイダーへ移行してください。
|
||||
|
||||
Unix-local から Docker に切り替えるには、エージェント定義は同じままにして、実行設定だけを変更します。
|
||||
|
||||
```python
|
||||
from docker import from_env as docker_from_env
|
||||
|
||||
from agents.run import RunConfig
|
||||
from agents.sandbox import SandboxRunConfig
|
||||
from agents.sandbox.sandboxes.docker import DockerSandboxClient, DockerSandboxClientOptions
|
||||
|
||||
run_config = RunConfig(
|
||||
sandbox=SandboxRunConfig(
|
||||
client=DockerSandboxClient(docker_from_env()),
|
||||
options=DockerSandboxClientOptions(image="python:3.14-slim"),
|
||||
),
|
||||
)
|
||||
```
|
||||
|
||||
コンテナ分離またはイメージの同等性が必要な場合に使用してください。[examples/sandbox/docker/docker_runner.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/docker/docker_runner.py) を参照してください。
|
||||
|
||||
## マウントとリモートストレージ
|
||||
|
||||
マウントエントリーはどのストレージを公開するかを表し、マウント戦略はサンドボックスバックエンドがそのストレージをどのようにアタッチするかを表します。組み込みのマウントエントリーと汎用戦略は `agents.sandbox.entries` からインポートします。ホスト型プロバイダーの戦略は `agents.extensions.sandbox` またはプロバイダー固有の拡張パッケージから利用できます。
|
||||
|
||||
一般的なマウントオプション:
|
||||
|
||||
- `mount_path`: ストレージがサンドボックス内で表示される場所です。相対パスはマニフェストルート配下で解決され、絶対パスはそのまま使用されます。
|
||||
- `read_only`: 既定は `True` です。サンドボックスがマウントされたストレージへ書き戻す必要がある場合にのみ `False` に設定してください。
|
||||
- `mount_strategy`: 必須です。マウントエントリーとサンドボックスバックエンドの両方に合う戦略を使用してください。
|
||||
|
||||
マウントは一時的なワークスペースエントリーとして扱われます。スナップショットと永続化のフローでは、マウントされたリモートストレージを保存済みワークスペースへコピーするのではなく、マウントされたパスをデタッチするかスキップします。
|
||||
|
||||
汎用ローカル / コンテナ戦略:
|
||||
|
||||
<div class="sandbox-nowrap-first-column-table" markdown="1">
|
||||
|
||||
| 戦略またはパターン | 使用する場面 | 備考 |
|
||||
| --- | --- | --- |
|
||||
| `InContainerMountStrategy(pattern=RcloneMountPattern(...))` | サンドボックスイメージで `rclone` を実行できる場合。 | S3、GCS、R2、Azure Blob、Box をサポートします。`RcloneMountPattern` は `fuse` モードまたは `nfs` モードで実行できます。 |
|
||||
| `InContainerMountStrategy(pattern=MountpointMountPattern(...))` | イメージに `mount-s3` があり、Mountpoint スタイルの S3 または S3 互換アクセスが必要な場合。 | `S3Mount` と `GCSMount` をサポートします。 |
|
||||
| `InContainerMountStrategy(pattern=FuseMountPattern(...))` | イメージに `blobfuse2` があり、FUSE サポートがある場合。 | `AzureBlobMount` をサポートします。 |
|
||||
| `InContainerMountStrategy(pattern=S3FilesMountPattern(...))` | イメージに `mount.s3files` があり、既存の S3 Files マウントターゲットに到達できる場合。 | `S3FilesMount` をサポートします。 |
|
||||
| `DockerVolumeMountStrategy(driver=...)` | Docker がコンテナ起動前にボリュームドライバー対応のマウントをアタッチする必要がある場合。 | Docker のみです。`rclone` は S3、GCS、R2、Azure Blob、Box をサポートし、`mountpoint` は S3 と GCS もサポートします。 |
|
||||
|
||||
</div>
|
||||
|
||||
## サポートされるホスト型プラットフォーム
|
||||
|
||||
ホスト型環境が必要な場合、通常は同じ `SandboxAgent` 定義をそのまま引き継ぎ、[`SandboxRunConfig`][agents.run_config.SandboxRunConfig] でサンドボックスクライアントだけを変更します。
|
||||
|
||||
このリポジトリのチェックアウトではなく公開されている SDK を使用している場合は、対応するパッケージ extra を通じてサンドボックスクライアントの依存関係をインストールしてください。
|
||||
|
||||
プロバイダー固有のセットアップメモと、チェックイン済みの拡張コード例へのリンクについては、[examples/sandbox/extensions/README.md](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/extensions/README.md) を参照してください。
|
||||
|
||||
<div class="sandbox-nowrap-first-column-table" markdown="1">
|
||||
|
||||
| クライアント | インストール | 例 |
|
||||
| --- | --- | --- |
|
||||
| `BlaxelSandboxClient` | `openai-agents[blaxel]` | [Blaxel ランナー](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/extensions/blaxel_runner.py) |
|
||||
| `CloudflareSandboxClient` | `openai-agents[cloudflare]` | [Cloudflare ランナー](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/extensions/cloudflare_runner.py) |
|
||||
| `DaytonaSandboxClient` | `openai-agents[daytona]` | [Daytona ランナー](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/extensions/daytona/daytona_runner.py) |
|
||||
| `E2BSandboxClient` | `openai-agents[e2b]` | [E2B ランナー](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/extensions/e2b_runner.py) |
|
||||
| `ModalSandboxClient` | `openai-agents[modal]` | [Modal ランナー](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/extensions/modal_runner.py) |
|
||||
| `RunloopSandboxClient` | `openai-agents[runloop]` | [Runloop ランナー](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/extensions/runloop/runner.py) |
|
||||
| `VercelSandboxClient` | `openai-agents[vercel]` | [Vercel ランナー](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/extensions/vercel_runner.py) |
|
||||
|
||||
</div>
|
||||
|
||||
ホスト型サンドボックスクライアントは、プロバイダー固有のマウント戦略を公開します。ストレージプロバイダーに最も合うバックエンドとマウント戦略を選択してください。
|
||||
|
||||
<div class="sandbox-nowrap-first-column-table" markdown="1">
|
||||
|
||||
| バックエンド | マウントに関する注記 |
|
||||
| --- | --- |
|
||||
| Docker | `InContainerMountStrategy` や `DockerVolumeMountStrategy` などのローカル戦略で、`S3Mount`、`GCSMount`、`R2Mount`、`AzureBlobMount`、`BoxMount`、`S3FilesMount` をサポートします。 |
|
||||
| `ModalSandboxClient` | `S3Mount`、`R2Mount`、HMAC 認証済みの `GCSMount` で、`ModalCloudBucketMountStrategy` による Modal のクラウドバケットマウントをサポートします。インライン認証情報、または名前付きの Modal Secret を使用できます。 |
|
||||
| `CloudflareSandboxClient` | `S3Mount`、`R2Mount`、HMAC 認証済みの `GCSMount` で、`CloudflareBucketMountStrategy` による Cloudflare バケットマウントをサポートします。 |
|
||||
| `BlaxelSandboxClient` | `S3Mount`、`R2Mount`、`GCSMount` で、`BlaxelCloudBucketMountStrategy` によるクラウドバケットマウントをサポートします。`agents.extensions.sandbox.blaxel` の `BlaxelDriveMount` と `BlaxelDriveMountStrategy` による永続的な Blaxel Drives もサポートします。 |
|
||||
| `DaytonaSandboxClient` | `DaytonaCloudBucketMountStrategy` による `rclone` ベースのクラウドストレージマウントをサポートします。`S3Mount`、`GCSMount`、`R2Mount`、`AzureBlobMount`、`BoxMount` と組み合わせて使用してください。 |
|
||||
| `E2BSandboxClient` | `E2BCloudBucketMountStrategy` による `rclone` ベースのクラウドストレージマウントをサポートします。`S3Mount`、`GCSMount`、`R2Mount`、`AzureBlobMount`、`BoxMount` と組み合わせて使用してください。 |
|
||||
| `RunloopSandboxClient` | `RunloopCloudBucketMountStrategy` による `rclone` ベースのクラウドストレージマウントをサポートします。`S3Mount`、`GCSMount`、`R2Mount`、`AzureBlobMount`、`BoxMount` と組み合わせて使用してください。 |
|
||||
| `VercelSandboxClient` | 現時点ではホスト型固有のマウント戦略は公開されていません。代わりにマニフェストファイル、リポジトリ、またはその他のワークスペース入力を使用してください。 |
|
||||
|
||||
</div>
|
||||
|
||||
以下の表は、各バックエンドが直接マウントできるリモートストレージエントリーをまとめたものです。
|
||||
|
||||
<div class="sandbox-nowrap-first-column-table" markdown="1">
|
||||
|
||||
| バックエンド | AWS S3 | Cloudflare R2 | GCS | Azure Blob Storage | Box | S3 Files |
|
||||
| --- | --- | --- | --- | --- | --- | --- |
|
||||
| Docker | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
|
||||
| `ModalSandboxClient` | ✓ | ✓ | ✓ | - | - | - |
|
||||
| `CloudflareSandboxClient` | ✓ | ✓ | ✓ | - | - | - |
|
||||
| `BlaxelSandboxClient` | ✓ | ✓ | ✓ | - | - | - |
|
||||
| `DaytonaSandboxClient` | ✓ | ✓ | ✓ | ✓ | ✓ | - |
|
||||
| `E2BSandboxClient` | ✓ | ✓ | ✓ | ✓ | ✓ | - |
|
||||
| `RunloopSandboxClient` | ✓ | ✓ | ✓ | ✓ | ✓ | - |
|
||||
| `VercelSandboxClient` | - | - | - | - | - | - |
|
||||
|
||||
</div>
|
||||
|
||||
実行可能なコード例をさらに見るには、ローカル、コーディング、メモリ、ハンドオフ、エージェント合成パターンについては [examples/sandbox/](https://github.com/openai/openai-agents-python/tree/main/examples/sandbox) を、ホスト型サンドボックスクライアントについては [examples/sandbox/extensions/](https://github.com/openai/openai-agents-python/tree/main/examples/sandbox/extensions) を参照してください。
|
||||
@@ -0,0 +1,863 @@
|
||||
---
|
||||
search:
|
||||
exclude: true
|
||||
---
|
||||
# 概念
|
||||
|
||||
!!! warning "ベータ機能"
|
||||
|
||||
サンドボックスエージェントはベータ版です。一般提供前に API の詳細、デフォルト、サポートされる機能が変更される可能性があります。また、時間の経過とともにより高度な機能が追加される予定です。
|
||||
|
||||
現代的なエージェントは、ファイルシステム上の実ファイルを操作できるときに最も効果的に動作します。**サンドボックスエージェント** は、専用ツールとシェルコマンドを使用して、大規模なドキュメントセットの検索や操作、ファイルの編集、成果物の生成、コマンドの実行を行えます。サンドボックスは、エージェントがユーザーの代わりに作業するために利用できる永続的なワークスペースをモデルに提供します。Agents SDK のサンドボックスエージェントは、サンドボックス環境と組み合わせたエージェントを簡単に実行できるようにし、ファイルシステム上に適切なファイルを配置し、サンドボックスをオーケストレーションして、タスクを大規模に開始、停止、再開しやすくします。
|
||||
|
||||
エージェントが必要とするデータを中心にワークスペースを定義します。ワークスペースは、GitHub リポジトリ、ローカルファイルとディレクトリ、合成タスクファイル、S3 や Azure Blob Storage などのリモートファイルシステム、その他提供するサンドボックス入力から開始できます。
|
||||
|
||||
<div class="sandbox-harness-image" markdown="1">
|
||||
|
||||

|
||||
|
||||
</div>
|
||||
|
||||
`SandboxAgent` は引き続き `Agent` です。`instructions`、`prompt`、`tools`、`handoffs`、`mcp_servers`、`model_settings`、`output_type`、ガードレール、フックなど、通常のエージェントのインターフェイスを保持し、通常の `Runner` API を通じて実行されます。変わるのは実行境界です。
|
||||
|
||||
- `SandboxAgent` はエージェント自体を定義します。通常のエージェント設定に加えて、`default_manifest`、`base_instructions`、`run_as` などのサンドボックス固有のデフォルト、ファイルシステムツール、シェルアクセス、スキル、メモリ、コンパクションなどの機能を含みます。
|
||||
- `Manifest` は、ファイル、リポジトリ、マウント、環境など、新規サンドボックスワークスペースの望ましい初期内容とレイアウトを宣言します。
|
||||
- サンドボックスセッションは、コマンドが実行されファイルが変更される、ライブの隔離環境です。
|
||||
- [`SandboxRunConfig`][agents.run_config.SandboxRunConfig] は、実行がサンドボックスセッションをどのように取得するかを決定します。たとえば、直接注入する、シリアライズ済みのサンドボックスセッション状態から再接続する、またはサンドボックスクライアントを通じて新規サンドボックスセッションを作成する、といった方法があります。
|
||||
- 保存済みのサンドボックス状態とスナップショットにより、後続の実行は以前の作業に再接続したり、保存済みの内容から新規サンドボックスセッションを初期化したりできます。
|
||||
|
||||
`Manifest` は新規セッションのワークスペース契約であり、すべてのライブサンドボックスに対する完全な唯一の情報源ではありません。実行の有効なワークスペースは、再利用されたサンドボックスセッション、シリアライズ済みのサンドボックスセッション状態、または実行時に選択されたスナップショットから取得される場合があります。
|
||||
|
||||
このページ全体で、「サンドボックスセッション」とはサンドボックスクライアントによって管理されるライブ実行環境を意味します。これは、[Sessions](../sessions/index.md) で説明されている SDK の会話型 [`Session`][agents.memory.session.Session] インターフェイスとは異なります。
|
||||
|
||||
外側のランタイムは引き続き、承認、トレーシング、ハンドオフ、再開の記録管理を所有します。サンドボックスセッションは、コマンド、ファイル変更、環境隔離を所有します。この分担は、このモデルの中核です。
|
||||
|
||||
### 各要素の関係
|
||||
|
||||
サンドボックス実行は、エージェント定義と実行ごとのサンドボックス設定を組み合わせます。Runner はエージェントを準備し、ライブのサンドボックスセッションにバインドし、後続の実行のために状態を保存できます。
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
agent["SandboxAgent<br/><small>full Agent + sandbox defaults</small>"]
|
||||
config["SandboxRunConfig<br/><small>client / session / resume inputs</small>"]
|
||||
runner["Runner<br/><small>prepare instructions<br/>bind capability tools</small>"]
|
||||
sandbox["sandbox session<br/><small>workspace where commands run<br/>and files change</small>"]
|
||||
saved["saved state / snapshot<br/><small>for resume or fresh-start later</small>"]
|
||||
|
||||
agent --> runner
|
||||
config --> runner
|
||||
runner --> sandbox
|
||||
sandbox --> saved
|
||||
```
|
||||
|
||||
サンドボックス固有のデフォルトは `SandboxAgent` に保持します。実行ごとのサンドボックスセッションの選択は `SandboxRunConfig` に保持します。
|
||||
|
||||
ライフサイクルは 3 つのフェーズで考えてください。
|
||||
|
||||
1. `SandboxAgent`、`Manifest`、機能を使って、エージェントと新規ワークスペース契約を定義します。
|
||||
2. サンドボックスセッションを注入、再開、または作成する `SandboxRunConfig` を `Runner` に渡して、実行を行います。
|
||||
3. Runner 管理の `RunState`、明示的なサンドボックス `session_state`、または保存済みワークスペーススナップショットから、後で継続します。
|
||||
|
||||
シェルアクセスが時々使うツールの 1 つにすぎない場合は、[ツールガイド](../tools.md) のホスト型シェルから始めてください。ワークスペースの隔離、サンドボックスクライアントの選択、またはサンドボックスセッションの再開動作が設計の一部である場合は、サンドボックスエージェントを選択してください。
|
||||
|
||||
## 使用すべき場面
|
||||
|
||||
サンドボックスエージェントは、ワークスペース中心のワークフローに適しています。例:
|
||||
|
||||
- コーディングとデバッグ。たとえば、GitHub リポジトリ内の issue レポートに対する自動修正をオーケストレーションし、対象を絞ったテストを実行する場合
|
||||
- ドキュメント処理と編集。たとえば、ユーザーの財務ドキュメントから情報を抽出し、完成済みの税務フォーム草案を作成する場合
|
||||
- ファイルに基づくレビューまたは分析。たとえば、回答前にオンボーディング資料、生成されたレポート、成果物バンドルを確認する場合
|
||||
- 隔離されたマルチエージェントパターン。たとえば、各レビュアーまたはコーディングサブエージェントに独自のワークスペースを与える場合
|
||||
- 複数ステップのワークスペースタスク。たとえば、ある実行でバグを修正し、後で回帰テストを追加する場合、またはスナップショットやサンドボックスセッション状態から再開する場合
|
||||
|
||||
ファイルや生きたファイルシステムへのアクセスが不要な場合は、`Agent` を使い続けてください。シェルアクセスが時々使う機能の 1 つにすぎない場合は、ホスト型シェルを追加してください。ワークスペース境界自体が機能の一部である場合は、サンドボックスエージェントを使用してください。
|
||||
|
||||
## サンドボックスクライアントの選択
|
||||
|
||||
ローカル開発では `UnixLocalSandboxClient` から始めてください。コンテナ隔離やイメージの一致性が必要になったら `DockerSandboxClient` に移行してください。プロバイダー管理の実行が必要になったら、ホスト型プロバイダーに移行してください。
|
||||
|
||||
ほとんどの場合、`SandboxAgent` 定義は同じままで、サンドボックスクライアントとそのオプションを [`SandboxRunConfig`][agents.run_config.SandboxRunConfig] で変更します。ローカル、Docker、ホスト型、リモートマウントのオプションについては、[サンドボックスクライアント](clients.md) を参照してください。
|
||||
|
||||
## 主要な構成要素
|
||||
|
||||
<div class="sandbox-nowrap-first-column-table" markdown="1">
|
||||
|
||||
| レイヤー | 主な SDK 構成要素 | 答える問い |
|
||||
| --- | --- | --- |
|
||||
| エージェント定義 | `SandboxAgent`, `Manifest`, capabilities | どのエージェントを実行し、新規セッションのワークスペース契約を何から開始すべきですか? |
|
||||
| サンドボックス実行 | `SandboxRunConfig`、サンドボックスクライアント、ライブサンドボックスセッション | この実行はライブサンドボックスセッションをどのように取得し、作業はどこで実行されますか? |
|
||||
| 保存済みサンドボックス状態 | `RunState` サンドボックスペイロード、`session_state`、スナップショット | このワークフローは以前のサンドボックス作業にどのように再接続するか、または保存済み内容から新規サンドボックスセッションをどのように初期化しますか? |
|
||||
|
||||
</div>
|
||||
|
||||
主な SDK 構成要素は、これらのレイヤーに次のように対応します。
|
||||
|
||||
<div class="sandbox-nowrap-first-column-table" markdown="1">
|
||||
|
||||
| 構成要素 | 管理するもの | 問うべきこと |
|
||||
| --- | --- | --- |
|
||||
| [`SandboxAgent`][agents.sandbox.sandbox_agent.SandboxAgent] | エージェント定義 | このエージェントは何を行うべきで、どのデフォルトを一緒に持ち運ぶべきですか? |
|
||||
| [`Manifest`][agents.sandbox.manifest.Manifest] | 新規セッションのワークスペースファイルとフォルダー | 実行開始時に、ファイルシステム上にどのファイルとフォルダーが存在すべきですか? |
|
||||
| [`Capability`][agents.sandbox.capabilities.capability.Capability] | サンドボックスネイティブな挙動 | どのツール、指示の断片、またはランタイム挙動をこのエージェントに付与すべきですか? |
|
||||
| [`SandboxRunConfig`][agents.run_config.SandboxRunConfig] | 実行ごとのサンドボックスクライアントとサンドボックスセッションの取得元 | この実行ではサンドボックスセッションを注入、再開、または作成すべきですか? |
|
||||
| [`RunState`][agents.run_state.RunState] | Runner が管理する保存済みサンドボックス状態 | 以前の Runner 管理ワークフローを再開し、そのサンドボックス状態を自動的に引き継いでいますか? |
|
||||
| [`SandboxRunConfig.session_state`][agents.run_config.SandboxRunConfig.session_state] | 明示的なシリアライズ済みサンドボックスセッション状態 | `RunState` の外部で既にシリアライズしたサンドボックス状態から再開したいですか? |
|
||||
| [`SandboxRunConfig.snapshot`][agents.run_config.SandboxRunConfig.snapshot] | 新規サンドボックスセッション用の保存済みワークスペース内容 | 新規サンドボックスセッションを保存済みファイルや成果物から開始すべきですか? |
|
||||
|
||||
</div>
|
||||
|
||||
実践的な設計順序は次のとおりです。
|
||||
|
||||
1. `Manifest` で新規セッションのワークスペース契約を定義します。
|
||||
2. `SandboxAgent` でエージェントを定義します。
|
||||
3. 組み込みまたはカスタムの機能を追加します。
|
||||
4. 各実行がサンドボックスセッションをどのように取得するかを `RunConfig(sandbox=SandboxRunConfig(...))` で決定します。
|
||||
|
||||
## サンドボックス実行の準備
|
||||
|
||||
実行時に、Runner はその定義を具体的なサンドボックス backed の実行に変換します。
|
||||
|
||||
1. `SandboxRunConfig` からサンドボックスセッションを解決します。
|
||||
`session=...` を渡すと、そのライブサンドボックスセッションを再利用します。
|
||||
それ以外の場合は、`client=...` を使用して作成または再開します。
|
||||
2. 実行の有効なワークスペース入力を決定します。
|
||||
実行がサンドボックスセッションを注入または再開する場合、その既存のサンドボックス状態が優先されます。
|
||||
それ以外の場合、Runner は 1 回限りのマニフェストオーバーライド、または `agent.default_manifest` から開始します。
|
||||
そのため、`Manifest` だけでは、すべての実行における最終的なライブワークスペースは定義されません。
|
||||
3. 機能に、結果として得られたマニフェストを処理させます。
|
||||
これにより、最終的なエージェントを準備する前に、機能がファイル、マウント、またはその他のワークスペーススコープの挙動を追加できます。
|
||||
4. 固定された順序で最終的な instructions を構築します。
|
||||
SDK のデフォルトサンドボックスプロンプト、または明示的にオーバーライドした場合は `base_instructions`、次に `instructions`、次に機能の指示断片、次にリモートマウントのポリシーテキスト、最後にレンダリングされたファイルシステムツリーです。
|
||||
5. 機能ツールをライブサンドボックスセッションにバインドし、通常の `Runner` API を通じて準備済みエージェントを実行します。
|
||||
|
||||
サンドボックス化によって、ターンの意味は変わりません。ターンは引き続きモデルステップであり、単一のシェルコマンドやサンドボックスアクションではありません。サンドボックス側の操作とターンの間に固定の 1:1 対応はありません。一部の作業はサンドボックス実行レイヤー内に留まる一方、別のアクションはツール結果、承認、または次のモデルステップを必要とするその他の状態を返す場合があります。実践上の規則として、サンドボックス作業の後にエージェントランタイムが別のモデル応答を必要とする場合にのみ、もう 1 ターンが消費されます。
|
||||
|
||||
これらの準備手順があるため、`SandboxAgent` を設計するときに考えるべき主なサンドボックス固有オプションは、`default_manifest`、`instructions`、`base_instructions`、`capabilities`、`run_as` です。
|
||||
|
||||
## `SandboxAgent` のオプション
|
||||
|
||||
通常の `Agent` フィールドに加えて用意されている、サンドボックス固有のオプションは次のとおりです。
|
||||
|
||||
<div class="sandbox-nowrap-first-column-table" markdown="1">
|
||||
|
||||
| オプション | 主な用途 |
|
||||
| --- | --- |
|
||||
| `default_manifest` | Runner が作成する新規サンドボックスセッションのデフォルトワークスペース。 |
|
||||
| `instructions` | SDK サンドボックスプロンプトの後に追加される、追加の役割、ワークフロー、成功基準。 |
|
||||
| `base_instructions` | SDK サンドボックスプロンプトを置き換える高度なオーバーライド手段。 |
|
||||
| `capabilities` | このエージェントと一緒に持ち運ぶべきサンドボックスネイティブなツールと挙動。 |
|
||||
| `run_as` | シェルコマンド、ファイル読み取り、パッチなど、モデル向けサンドボックスツールのユーザー ID。 |
|
||||
|
||||
</div>
|
||||
|
||||
サンドボックスクライアントの選択、サンドボックスセッションの再利用、マニフェストオーバーライド、スナップショット選択は、エージェントではなく [`SandboxRunConfig`][agents.run_config.SandboxRunConfig] に属します。
|
||||
|
||||
### `default_manifest`
|
||||
|
||||
`default_manifest` は、Runner がこのエージェント用に新規サンドボックスセッションを作成するときに使用されるデフォルトの [`Manifest`][agents.sandbox.manifest.Manifest] です。エージェントが通常開始すべきファイル、リポジトリ、補助資料、出力ディレクトリ、マウントに使用します。
|
||||
|
||||
これはデフォルトにすぎません。実行は `SandboxRunConfig(manifest=...)` でこれをオーバーライドできます。また、再利用または再開されたサンドボックスセッションは、既存のワークスペース状態を保持します。
|
||||
|
||||
### `instructions` と `base_instructions`
|
||||
|
||||
異なるプロンプトでも維持すべき短いルールには `instructions` を使用します。`SandboxAgent` では、これらの instructions は SDK のサンドボックスベースプロンプトの後に追加されるため、組み込みのサンドボックスガイダンスを保持しつつ、独自の役割、ワークフロー、成功基準を追加できます。
|
||||
|
||||
SDK サンドボックスベースプロンプトを置き換えたい場合にのみ、`base_instructions` を使用してください。ほとんどのエージェントでは設定すべきではありません。
|
||||
|
||||
<div class="sandbox-nowrap-first-column-table" markdown="1">
|
||||
|
||||
| 配置先... | 用途 | 例 |
|
||||
| --- | --- | --- |
|
||||
| `instructions` | エージェントの安定した役割、ワークフロールール、成功基準。 | 「オンボーディングドキュメントを調査してから、ハンドオフします。」、「最終ファイルを `output/` に書き込みます。」 |
|
||||
| `base_instructions` | SDK サンドボックスベースプロンプトの完全な置き換え。 | カスタムの低レベルサンドボックスラッパープロンプト。 |
|
||||
| ユーザープロンプト | この実行の 1 回限りのリクエスト。 | 「このワークスペースを要約してください。」 |
|
||||
| マニフェスト内のワークスペースファイル | より長いタスク仕様、リポジトリローカルの指示、または範囲が限定された参照資料。 | `repo/task.md`、ドキュメントバンドル、サンプルパケット。 |
|
||||
|
||||
</div>
|
||||
|
||||
`instructions` の適切な用途には、次のようなものがあります。
|
||||
|
||||
- [examples/sandbox/unix_local_pty.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/unix_local_pty.py) は、PTY 状態が重要な場合に、エージェントを 1 つの対話型プロセス内に留めます。
|
||||
- [examples/sandbox/handoffs.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/handoffs.py) は、サンドボックスレビュアーが検査後にユーザーへ直接回答することを禁止します。
|
||||
- [examples/sandbox/tax_prep.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/tax_prep.py) は、最終的に記入済みのファイルが実際に `output/` に配置されることを要求します。
|
||||
- [examples/sandbox/docs/coding_task.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/docs/coding_task.py) は、正確な検証コマンドを固定し、ワークスペースルート相対のパッチパスを明確にします。
|
||||
|
||||
ユーザーの 1 回限りのタスクを `instructions` にコピーすること、マニフェストに置くべき長い参照資料を埋め込むこと、組み込み機能が既に注入するツールドキュメントを繰り返すこと、または実行時にモデルが必要としないローカルインストールメモを混ぜることは避けてください。
|
||||
|
||||
`instructions` を省略しても、SDK はデフォルトのサンドボックスプロンプトを含めます。低レベルのラッパーにはそれで十分ですが、ほとんどのユーザー向けエージェントでは、明示的な `instructions` を提供すべきです。
|
||||
|
||||
### `capabilities`
|
||||
|
||||
機能は、サンドボックスネイティブな挙動を `SandboxAgent` に付与します。実行開始前にワークスペースを形作り、サンドボックス固有の指示を追加し、ライブサンドボックスセッションにバインドされるツールを公開し、そのエージェントのモデル挙動や入力処理を調整できます。
|
||||
|
||||
組み込み機能には次のものがあります。
|
||||
|
||||
<div class="sandbox-nowrap-first-column-table" markdown="1">
|
||||
|
||||
| 機能 | 追加する場面 | 注意 |
|
||||
| --- | --- | --- |
|
||||
| `Shell` | エージェントがシェルアクセスを必要とする場合。 | `exec_command` を追加し、サンドボックスクライアントが PTY インタラクションをサポートする場合は `write_stdin` も追加します。 |
|
||||
| `Filesystem` | エージェントがファイルを編集する、またはローカル画像を検査する必要がある場合。 | `apply_patch` と `view_image` を追加します。パッチパスはワークスペースルート相対です。 |
|
||||
| `Skills` | サンドボックス内でスキル検出とマテリアライズを行いたい場合。 | `.agents` または `.agents/skills` を手動でマウントするよりもこちらを優先してください。`Skills` はスキルのインデックス化とサンドボックスへのマテリアライズを行います。 |
|
||||
| `Memory` | 後続の実行でメモリ成果物を読み取る、または生成する必要がある場合。 | `Shell` が必要です。ライブ更新には `Filesystem` も必要です。 |
|
||||
| `Compaction` | 長時間実行されるフローで、コンパクション項目の後にコンテキストのトリミングが必要な場合。 | モデルサンプリングと入力処理を調整します。 |
|
||||
|
||||
</div>
|
||||
|
||||
デフォルトでは、`SandboxAgent.capabilities` は `Capabilities.default()` を使用します。これには `Filesystem()`、`Shell()`、`Compaction()` が含まれます。`capabilities=[...]` を渡すと、そのリストがデフォルトを置き換えるため、引き続き必要なデフォルト機能を含めてください。
|
||||
|
||||
スキルについては、どのようにマテリアライズしたいかに基づいてソースを選択してください。
|
||||
|
||||
- `Skills(lazy_from=LocalDirLazySkillSource(...))` は、大きめのローカルスキルディレクトリに適したデフォルトです。モデルがまずインデックスを発見し、必要なものだけを読み込めるためです。
|
||||
- `LocalDirLazySkillSource(source=LocalDir(src=...))` は、SDK プロセスが実行されているファイルシステムから読み取ります。サンドボックスイメージまたはワークスペース内にのみ存在するパスではなく、元のホスト側スキルディレクトリを渡してください。
|
||||
- `Skills(from_=LocalDir(src=...))` は、事前にステージングしたい小さなローカルバンドルに適しています。
|
||||
- `Skills(from_=GitRepo(repo=..., ref=...))` は、スキル自体をリポジトリから取得すべき場合に適しています。
|
||||
|
||||
`LocalDir.src` は SDK ホスト上のソースパスです。`skills_path` は、`load_skill` が呼び出されたときにスキルがステージングされる、サンドボックスワークスペース内の相対宛先パスです。
|
||||
|
||||
スキルが既に `.agents/skills/<name>/SKILL.md` のような場所にディスク上で存在する場合は、そのソースルートを `LocalDir(...)` に指定し、それでも `Skills(...)` を使用して公開してください。別のサンドボックス内レイアウトに依存する既存のワークスペース契約がない限り、デフォルトの `skills_path=".agents"` を維持してください。
|
||||
|
||||
適合する場合は、組み込み機能を優先してください。組み込みでカバーされないサンドボックス固有のツールや指示面が必要な場合にのみ、カスタム機能を作成してください。
|
||||
|
||||
## 概念
|
||||
|
||||
### マニフェスト
|
||||
|
||||
[`Manifest`][agents.sandbox.manifest.Manifest] は、新規サンドボックスセッションのワークスペースを記述します。ワークスペースの `root` を設定し、ファイルとディレクトリを宣言し、ローカルファイルをコピーし、Git リポジトリをクローンし、リモートストレージマウントを接続し、環境変数を設定し、ユーザーやグループを定義し、ワークスペース外の特定の絶対パスへのアクセスを許可できます。
|
||||
|
||||
マニフェストエントリのパスはワークスペース相対です。絶対パスにすることや、`..` でワークスペース外へ抜けることはできません。これにより、ワークスペース契約はローカル、Docker、ホスト型クライアント間で移植可能になります。
|
||||
|
||||
作業開始前にエージェントが必要とする資料には、マニフェストエントリを使用します。
|
||||
|
||||
<div class="sandbox-nowrap-first-column-table" markdown="1">
|
||||
|
||||
| マニフェストエントリ | 用途 |
|
||||
| --- | --- |
|
||||
| `File`, `Dir` | 小さな合成入力、補助ファイル、または出力ディレクトリ。 |
|
||||
| `LocalFile`, `LocalDir` | サンドボックスにマテリアライズすべきホストファイルまたはディレクトリ。 |
|
||||
| `GitRepo` | ワークスペースに取得すべきリポジトリ。 |
|
||||
| `S3Mount`, `GCSMount`, `R2Mount`, `AzureBlobMount`, `BoxMount`, `S3FilesMount` などのマウント | サンドボックス内に表示すべき外部ストレージ。 |
|
||||
|
||||
</div>
|
||||
|
||||
`Dir` は、合成子要素から、または出力場所として、サンドボックスワークスペース内にディレクトリを作成します。ホストファイルシステムから読み取るわけではありません。既存のホストディレクトリをサンドボックスワークスペースにコピーすべき場合は、`LocalDir` を使用してください。
|
||||
|
||||
`LocalFile.src` と `LocalDir.src` は、デフォルトでは SDK プロセスの作業ディレクトリを基準に解決されます。ソースは、`extra_path_grants` でカバーされていない限り、そのベースディレクトリの下に留まる必要があります。これにより、ローカルソースのマテリアライズは、サンドボックスマニフェストの他の部分と同じホストパスの信頼境界内に保たれます。
|
||||
|
||||
マウントエントリは公開するストレージを記述し、マウント戦略はサンドボックスバックエンドがそのストレージをどのように接続するかを記述します。マウントオプションとプロバイダーサポートについては、[サンドボックスクライアント](clients.md#mounts-and-remote-storage) を参照してください。
|
||||
|
||||
適切なマニフェスト設計では通常、ワークスペース契約を狭く保ち、長いタスク手順を `repo/task.md` などのワークスペースファイルに置き、`repo/task.md` や `output/report.md` などの相対ワークスペースパスを指示で使用します。エージェントが `Filesystem` 機能の `apply_patch` ツールでファイルを編集する場合、パッチパスはシェルの `workdir` ではなくサンドボックスワークスペースルートからの相対であることを忘れないでください。
|
||||
|
||||
エージェントがワークスペース外の具体的な絶対パスを必要とする場合、または SDK プロセスの作業ディレクトリ外にある信頼済みローカルソースをマニフェストがコピーする必要がある場合にのみ、`extra_path_grants` を使用してください。例として、一時的なツール出力用の `/tmp`、読み取り専用ランタイム用の `/opt/toolchain`、サンドボックスにマテリアライズすべき生成済みスキルディレクトリなどがあります。付与は、ローカルソースのマテリアライズ、SDK ファイル API、バックエンドがファイルシステムポリシーを適用できる場合のシェル実行に適用されます。
|
||||
|
||||
```python
|
||||
from agents.sandbox import Manifest, SandboxPathGrant
|
||||
|
||||
manifest = Manifest(
|
||||
extra_path_grants=(
|
||||
SandboxPathGrant(path="/tmp"),
|
||||
SandboxPathGrant(path="/opt/toolchain", read_only=True),
|
||||
),
|
||||
)
|
||||
```
|
||||
|
||||
`extra_path_grants` を含むマニフェストは、信頼済み設定として扱ってください。アプリケーションがそれらのホストパスを既に承認していない限り、モデル出力やその他の信頼できないペイロードから付与を読み込まないでください。
|
||||
|
||||
スナップショットと `persist_workspace()` は、引き続きワークスペースルートのみを含みます。追加で許可されたパスはランタイムアクセスであり、永続的なワークスペース状態ではありません。
|
||||
|
||||
### 権限
|
||||
|
||||
`Permissions` は、マニフェストエントリのファイルシステム権限を制御します。これはサンドボックスがマテリアライズするファイルに関するものであり、モデル権限、承認ポリシー、API 認証情報に関するものではありません。
|
||||
|
||||
デフォルトでは、マニフェストエントリは所有者が読み取り/書き込み/実行可能で、グループとその他のユーザーが読み取り/実行可能です。ステージングされたファイルをプライベート、読み取り専用、または実行可能にすべき場合は、これをオーバーライドしてください。
|
||||
|
||||
```python
|
||||
from agents.sandbox import FileMode, Permissions
|
||||
from agents.sandbox.entries import File
|
||||
|
||||
private_notes = File(
|
||||
text="internal notes",
|
||||
permissions=Permissions(
|
||||
owner=FileMode.READ | FileMode.WRITE,
|
||||
group=FileMode.NONE,
|
||||
other=FileMode.NONE,
|
||||
),
|
||||
)
|
||||
```
|
||||
|
||||
`Permissions` は、所有者、グループ、その他のビットを個別に保持し、さらにエントリがディレクトリかどうかも保持します。直接構築することも、`Permissions.from_str(...)` でモード文字列からパースすることも、`Permissions.from_mode(...)` で OS モードから導出することもできます。
|
||||
|
||||
ユーザーは、作業を実行できるサンドボックス ID です。その ID をサンドボックスに存在させたい場合は、マニフェストに `User` を追加し、シェルコマンド、ファイル読み取り、パッチなどのモデル向けサンドボックスツールをそのユーザーとして実行すべき場合は `SandboxAgent.run_as` を設定してください。`run_as` がマニフェスト内にまだ存在しないユーザーを指している場合、Runner は有効なマニフェストにそのユーザーを追加します。
|
||||
|
||||
```python
|
||||
from agents import Runner
|
||||
from agents.run import RunConfig
|
||||
from agents.sandbox import FileMode, Manifest, Permissions, SandboxAgent, SandboxRunConfig, User
|
||||
from agents.sandbox.entries import Dir, LocalDir
|
||||
from agents.sandbox.sandboxes.unix_local import UnixLocalSandboxClient
|
||||
|
||||
analyst = User(name="analyst")
|
||||
|
||||
agent = SandboxAgent(
|
||||
name="Dataroom analyst",
|
||||
instructions="Review the files in `dataroom/` and write findings to `output/`.",
|
||||
default_manifest=Manifest(
|
||||
# Declare the sandbox user so manifest entries can grant access to it.
|
||||
users=[analyst],
|
||||
entries={
|
||||
"dataroom": LocalDir(
|
||||
src="./dataroom",
|
||||
# Let the analyst traverse and read the mounted dataroom, but not edit it.
|
||||
group=analyst,
|
||||
permissions=Permissions(
|
||||
owner=FileMode.READ | FileMode.EXEC,
|
||||
group=FileMode.READ | FileMode.EXEC,
|
||||
other=FileMode.NONE,
|
||||
),
|
||||
),
|
||||
"output": Dir(
|
||||
# Give the analyst a writable scratch/output directory for artifacts.
|
||||
group=analyst,
|
||||
permissions=Permissions(
|
||||
owner=FileMode.ALL,
|
||||
group=FileMode.ALL,
|
||||
other=FileMode.NONE,
|
||||
),
|
||||
),
|
||||
},
|
||||
),
|
||||
# Run model-facing sandbox actions as this user, so those permissions apply.
|
||||
run_as=analyst,
|
||||
)
|
||||
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
"Summarize the contracts and call out renewal dates.",
|
||||
run_config=RunConfig(
|
||||
sandbox=SandboxRunConfig(client=UnixLocalSandboxClient()),
|
||||
),
|
||||
)
|
||||
```
|
||||
|
||||
ファイルレベルの共有ルールも必要な場合は、ユーザーとマニフェストグループ、エントリの `group` メタデータを組み合わせてください。`run_as` ユーザーは誰がサンドボックスネイティブなアクションを実行するかを制御し、`Permissions` はサンドボックスがワークスペースをマテリアライズした後に、そのユーザーがどのファイルを読み取り、書き込み、実行できるかを制御します。
|
||||
|
||||
### スナップショット仕様
|
||||
|
||||
`SnapshotSpec` は、新規サンドボックスセッションで保存済みワークスペース内容をどこから復元し、どこへ永続化して戻すかを指定します。これはサンドボックスワークスペースのスナップショットポリシーであり、`session_state` は特定のサンドボックスバックエンドを再開するためのシリアライズ済み接続状態です。
|
||||
|
||||
ローカルの永続スナップショットには `LocalSnapshotSpec` を使用し、アプリがリモートスナップショットクライアントを提供する場合は `RemoteSnapshotSpec` を使用します。ローカルスナップショットのセットアップが利用できない場合はフォールバックとして no-op スナップショットが使用され、高度な呼び出し元はワークスペーススナップショットの永続化を望まない場合に明示的に使用できます。
|
||||
|
||||
```python
|
||||
from pathlib import Path
|
||||
|
||||
from agents.run import RunConfig
|
||||
from agents.sandbox import LocalSnapshotSpec, SandboxRunConfig
|
||||
from agents.sandbox.sandboxes.unix_local import UnixLocalSandboxClient
|
||||
|
||||
run_config = RunConfig(
|
||||
sandbox=SandboxRunConfig(
|
||||
client=UnixLocalSandboxClient(),
|
||||
snapshot=LocalSnapshotSpec(base_path=Path("/tmp/my-sandbox-snapshots")),
|
||||
)
|
||||
)
|
||||
```
|
||||
|
||||
Runner が新規サンドボックスセッションを作成すると、サンドボックスクライアントはそのセッション用のスナップショットインスタンスを構築します。開始時に、スナップショットが復元可能であれば、実行が継続する前にサンドボックスは保存済みワークスペース内容を復元します。クリーンアップ時には、Runner 所有のサンドボックスセッションがワークスペースをアーカイブし、スナップショットを通じて永続化して戻します。
|
||||
|
||||
`snapshot` を省略すると、ランタイムは可能な場合にデフォルトのローカルスナップショット場所を使用しようとします。それをセットアップできない場合は、no-op スナップショットにフォールバックします。マウントされたパスと一時パスは、永続的なワークスペース内容としてスナップショットにコピーされません。
|
||||
|
||||
### サンドボックスのライフサイクル
|
||||
|
||||
ライフサイクルモードは 2 つあります。**SDK 所有** と **開発者所有** です。
|
||||
|
||||
<div class="sandbox-lifecycle-diagram" markdown="1">
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
participant App
|
||||
participant Runner
|
||||
participant Client
|
||||
participant Sandbox
|
||||
|
||||
App->>Runner: Runner.run(..., SandboxRunConfig(client=...))
|
||||
Runner->>Client: create or resume sandbox
|
||||
Client-->>Runner: sandbox session
|
||||
Runner->>Sandbox: start, run tools
|
||||
Runner->>Sandbox: stop and persist snapshot
|
||||
Runner->>Client: delete runner-owned resources
|
||||
|
||||
App->>Client: create(...)
|
||||
Client-->>App: sandbox session
|
||||
App->>Sandbox: async with sandbox
|
||||
App->>Runner: Runner.run(..., SandboxRunConfig(session=sandbox))
|
||||
Runner->>Sandbox: run tools
|
||||
App->>Sandbox: cleanup on context exit / aclose()
|
||||
```
|
||||
|
||||
</div>
|
||||
|
||||
サンドボックスが 1 回の実行の間だけ存続すればよい場合は、SDK 所有のライフサイクルを使用します。`client`、任意の `manifest`、任意の `snapshot`、クライアント `options` を渡します。Runner はサンドボックスを作成または再開し、開始し、エージェントを実行し、スナップショット backed のワークスペース状態を永続化し、サンドボックスをシャットダウンし、クライアントに Runner 所有リソースをクリーンアップさせます。
|
||||
|
||||
```python
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
"Inspect the workspace and summarize what changed.",
|
||||
run_config=RunConfig(
|
||||
sandbox=SandboxRunConfig(client=UnixLocalSandboxClient()),
|
||||
),
|
||||
)
|
||||
```
|
||||
|
||||
サンドボックスを先に作成したい場合、1 つのライブサンドボックスを複数の実行で再利用したい場合、実行後にファイルを検査したい場合、自分で作成したサンドボックス上でストリーミングしたい場合、またはクリーンアップのタイミングを厳密に決めたい場合は、開発者所有のライフサイクルを使用します。`session=...` を渡すと、Runner はそのライブサンドボックスを使用しますが、代わりに閉じることはありません。
|
||||
|
||||
```python
|
||||
sandbox = await client.create(manifest=agent.default_manifest)
|
||||
|
||||
async with sandbox:
|
||||
run_config = RunConfig(sandbox=SandboxRunConfig(session=sandbox))
|
||||
await Runner.run(agent, "Analyze the files.", run_config=run_config)
|
||||
await Runner.run(agent, "Write the final report.", run_config=run_config)
|
||||
```
|
||||
|
||||
通常はコンテキストマネージャーの形を使用します。エントリ時にサンドボックスを開始し、終了時にセッションのクリーンアップライフサイクルを実行します。アプリでコンテキストマネージャーを使用できない場合は、ライフサイクルメソッドを直接呼び出してください。
|
||||
|
||||
```python
|
||||
sandbox = await client.create(
|
||||
manifest=agent.default_manifest,
|
||||
snapshot=LocalSnapshotSpec(base_path=Path("/tmp/my-sandbox-snapshots")),
|
||||
)
|
||||
try:
|
||||
await sandbox.start()
|
||||
await Runner.run(
|
||||
agent,
|
||||
"Analyze the files.",
|
||||
run_config=RunConfig(sandbox=SandboxRunConfig(session=sandbox)),
|
||||
)
|
||||
# Persist a checkpoint of the live workspace before doing more work.
|
||||
# `aclose()` also calls `stop()`, so this is only needed for an explicit mid-lifecycle save.
|
||||
await sandbox.stop()
|
||||
finally:
|
||||
await sandbox.aclose()
|
||||
```
|
||||
|
||||
`stop()` はスナップショット backed のワークスペース内容だけを永続化し、サンドボックスを破棄しません。`aclose()` は完全なセッションクリーンアップパスです。停止前フックを実行し、`stop()` を呼び出し、サンドボックスリソースをシャットダウンし、セッションスコープの依存関係を閉じます。
|
||||
|
||||
## `SandboxRunConfig` のオプション
|
||||
|
||||
[`SandboxRunConfig`][agents.run_config.SandboxRunConfig] は、サンドボックスセッションがどこから来るか、および新規セッションをどのように初期化すべきかを決定する、実行ごとのオプションを保持します。
|
||||
|
||||
### サンドボックスの取得元
|
||||
|
||||
これらのオプションは、Runner がサンドボックスセッションを再利用、再開、または作成すべきかを決定します。
|
||||
|
||||
<div class="sandbox-nowrap-first-column-table" markdown="1">
|
||||
|
||||
| オプション | 使用する場面 | 注意 |
|
||||
| --- | --- | --- |
|
||||
| `client` | Runner にサンドボックスセッションの作成、再開、クリーンアップを任せたい場合。 | ライブサンドボックス `session` を提供しない限り必須です。 |
|
||||
| `session` | 既にライブサンドボックスセッションを自分で作成している場合。 | 呼び出し元がライフサイクルを所有します。Runner はそのライブサンドボックスセッションを再利用します。 |
|
||||
| `session_state` | シリアライズ済みのサンドボックスセッション状態はあるが、ライブサンドボックスセッションオブジェクトはない場合。 | `client` が必要です。Runner はその明示的な状態から所有セッションとして再開します。 |
|
||||
|
||||
</div>
|
||||
|
||||
実際には、Runner は次の順序でサンドボックスセッションを解決します。
|
||||
|
||||
1. `run_config.sandbox.session` を注入した場合、そのライブサンドボックスセッションが直接再利用されます。
|
||||
2. それ以外で、実行が `RunState` から再開される場合、保存されているサンドボックスセッション状態が再開されます。
|
||||
3. それ以外で、`run_config.sandbox.session_state` を渡した場合、Runner はその明示的なシリアライズ済みサンドボックスセッション状態から再開します。
|
||||
4. それ以外の場合、Runner は新規サンドボックスセッションを作成します。その新規セッションでは、提供されている場合は `run_config.sandbox.manifest` を使用し、そうでない場合は `agent.default_manifest` を使用します。
|
||||
|
||||
### 新規セッションの入力
|
||||
|
||||
これらのオプションは、Runner が新規サンドボックスセッションを作成する場合にのみ意味を持ちます。
|
||||
|
||||
<div class="sandbox-nowrap-first-column-table" markdown="1">
|
||||
|
||||
| オプション | 使用する場面 | 注意 |
|
||||
| --- | --- | --- |
|
||||
| `manifest` | 1 回限りの新規セッションワークスペースオーバーライドが必要な場合。 | 省略時は `agent.default_manifest` にフォールバックします。 |
|
||||
| `snapshot` | 新規サンドボックスセッションをスナップショットから初期化すべき場合。 | 再開に似たフローやリモートスナップショットクライアントに有用です。 |
|
||||
| `options` | サンドボックスクライアントが作成時オプションを必要とする場合。 | Docker イメージ、Modal アプリ名、E2B テンプレート、タイムアウト、および同様のクライアント固有設定で一般的です。 |
|
||||
|
||||
</div>
|
||||
|
||||
### マテリアライズ制御
|
||||
|
||||
`concurrency_limits` は、サンドボックスのマテリアライズ作業をどの程度並列に実行できるかを制御します。大きなマニフェストやローカルディレクトリコピーでより厳密なリソース制御が必要な場合は、`SandboxConcurrencyLimits(manifest_entries=..., local_dir_files=...)` を使用してください。いずれかの値を `None` に設定すると、その特定の制限を無効にできます。
|
||||
|
||||
`archive_limits` は、アーカイブ抽出に対する SDK 側のリソースチェックを制御します。SDK のデフォルトしきい値を有効にするには `archive_limits=SandboxArchiveLimits()` を設定します。アーカイブにより厳密なリソース制御が必要な場合は、`SandboxArchiveLimits(max_input_bytes=..., max_extracted_bytes=..., max_members=...)` などの明示的な値を渡してください。SDK アーカイブリソース制限なしのデフォルト動作を維持するには `archive_limits=None` のままにし、個別のフィールドだけを無効にするにはそのフィールドを `None` に設定します。
|
||||
|
||||
留意すべき点がいくつかあります。
|
||||
|
||||
- 新規セッション: `manifest=` と `snapshot=` は、Runner が新規サンドボックスセッションを作成する場合にのみ適用されます。
|
||||
- 再開とスナップショット: `session_state=` は以前にシリアライズされたサンドボックス状態に再接続します。一方、`snapshot=` は保存済みワークスペース内容から新規サンドボックスセッションを初期化します。
|
||||
- クライアント固有オプション: `options=` はサンドボックスクライアントに依存します。Docker や多くのホスト型クライアントでは必須です。
|
||||
- 注入されたライブセッション: 実行中のサンドボックス `session` を渡す場合、機能駆動のマニフェスト更新は互換性のある非マウントエントリを追加できます。`manifest.root`、`manifest.environment`、`manifest.users`、`manifest.groups` を変更すること、既存エントリを削除すること、エントリタイプを置き換えること、マウントエントリを追加または変更することはできません。
|
||||
- Runner API: `SandboxAgent` の実行は引き続き、通常の `Runner.run()`、`Runner.run_sync()`、`Runner.run_streamed()` API を使用します。
|
||||
|
||||
## 完全な例: コーディングタスク
|
||||
|
||||
このコーディングスタイルの例は、デフォルトの出発点として適しています。
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
from pathlib import Path
|
||||
|
||||
from agents import ModelSettings, Runner
|
||||
from agents.run import RunConfig
|
||||
from agents.sandbox import Manifest, SandboxAgent, SandboxRunConfig
|
||||
from agents.sandbox.capabilities import (
|
||||
Capabilities,
|
||||
LocalDirLazySkillSource,
|
||||
Skills,
|
||||
)
|
||||
from agents.sandbox.entries import LocalDir
|
||||
from agents.sandbox.sandboxes.unix_local import UnixLocalSandboxClient
|
||||
|
||||
EXAMPLE_DIR = Path(__file__).resolve().parent
|
||||
HOST_REPO_DIR = EXAMPLE_DIR / "repo"
|
||||
HOST_SKILLS_DIR = EXAMPLE_DIR / "skills"
|
||||
TARGET_TEST_CMD = "sh tests/test_credit_note.sh"
|
||||
|
||||
|
||||
def build_agent(model: str) -> SandboxAgent[None]:
|
||||
return SandboxAgent(
|
||||
name="Sandbox engineer",
|
||||
model=model,
|
||||
instructions=(
|
||||
"Inspect the repo, make the smallest correct change, run the most relevant checks, "
|
||||
"and summarize the file changes and risks. "
|
||||
"Read `repo/task.md` before editing files. Stay grounded in the repository, preserve "
|
||||
"existing behavior, and mention the exact verification command you ran. "
|
||||
"Use the `$credit-note-fixer` skill before editing files. If the repo lives under "
|
||||
"`repo/`, remember that `apply_patch` paths stay relative to the sandbox workspace "
|
||||
"root, so edits still target `repo/...`."
|
||||
),
|
||||
# Put repos and task files in the manifest.
|
||||
default_manifest=Manifest(
|
||||
entries={
|
||||
"repo": LocalDir(src=HOST_REPO_DIR),
|
||||
}
|
||||
),
|
||||
capabilities=Capabilities.default() + [
|
||||
Skills(
|
||||
lazy_from=LocalDirLazySkillSource(
|
||||
# This is a host path read by the SDK process.
|
||||
# Requested skills are copied into `skills_path` in the sandbox.
|
||||
source=LocalDir(src=HOST_SKILLS_DIR),
|
||||
)
|
||||
),
|
||||
],
|
||||
model_settings=ModelSettings(tool_choice="required"),
|
||||
)
|
||||
|
||||
|
||||
async def main(model: str, prompt: str) -> None:
|
||||
result = await Runner.run(
|
||||
build_agent(model),
|
||||
prompt,
|
||||
run_config=RunConfig(
|
||||
sandbox=SandboxRunConfig(client=UnixLocalSandboxClient()),
|
||||
workflow_name="Sandbox coding example",
|
||||
),
|
||||
)
|
||||
print(result.final_output)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
asyncio.run(
|
||||
main(
|
||||
model="gpt-5.5",
|
||||
prompt=(
|
||||
"Open `repo/task.md`, use the `$credit-note-fixer` skill, fix the bug, "
|
||||
f"run `{TARGET_TEST_CMD}`, and summarize the change."
|
||||
),
|
||||
)
|
||||
)
|
||||
```
|
||||
|
||||
[examples/sandbox/docs/coding_task.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/docs/coding_task.py) を参照してください。この例は、Unix ローカル実行間で決定論的に検証できるように、小さなシェルベースのリポジトリを使用しています。実際のタスクリポジトリは、もちろん Python、JavaScript、またはその他何でもかまいません。
|
||||
|
||||
## 一般的なパターン
|
||||
|
||||
上記の完全な例から始めてください。多くの場合、同じ `SandboxAgent` をそのまま維持し、サンドボックスクライアント、サンドボックスセッションの取得元、またはワークスペースの取得元だけを変更できます。
|
||||
|
||||
### サンドボックスクライアントの切り替え
|
||||
|
||||
エージェント定義は同じままにし、実行設定だけを変更します。コンテナ隔離やイメージの一致性が必要な場合は Docker を使用し、プロバイダー管理の実行が必要な場合はホスト型プロバイダーを使用します。例とプロバイダーオプションについては、[サンドボックスクライアント](clients.md) を参照してください。
|
||||
|
||||
### ワークスペースのオーバーライド
|
||||
|
||||
エージェント定義は同じままにし、新規セッションのマニフェストだけを差し替えます。
|
||||
|
||||
```python
|
||||
from agents.run import RunConfig
|
||||
from agents.sandbox import Manifest, SandboxRunConfig
|
||||
from agents.sandbox.entries import GitRepo
|
||||
from agents.sandbox.sandboxes.unix_local import UnixLocalSandboxClient
|
||||
|
||||
run_config = RunConfig(
|
||||
sandbox=SandboxRunConfig(
|
||||
client=UnixLocalSandboxClient(),
|
||||
manifest=Manifest(
|
||||
entries={
|
||||
"repo": GitRepo(repo="openai/openai-agents-python", ref="main"),
|
||||
}
|
||||
),
|
||||
),
|
||||
)
|
||||
```
|
||||
|
||||
同じエージェントの役割を、エージェントを再構築せずに異なるリポジトリ、パケット、またはタスクバンドルに対して実行すべき場合に使用します。上記の検証済みコーディング例では、1 回限りのオーバーライドではなく `default_manifest` で同じパターンを示しています。
|
||||
|
||||
### サンドボックスセッションの注入
|
||||
|
||||
明示的なライフサイクル制御、実行後の検査、または出力コピーが必要な場合は、ライブサンドボックスセッションを注入します。
|
||||
|
||||
```python
|
||||
from agents import Runner
|
||||
from agents.run import RunConfig
|
||||
from agents.sandbox import SandboxRunConfig
|
||||
from agents.sandbox.sandboxes.unix_local import UnixLocalSandboxClient
|
||||
|
||||
client = UnixLocalSandboxClient()
|
||||
sandbox = await client.create(manifest=agent.default_manifest)
|
||||
|
||||
async with sandbox:
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
prompt,
|
||||
run_config=RunConfig(
|
||||
sandbox=SandboxRunConfig(session=sandbox),
|
||||
),
|
||||
)
|
||||
```
|
||||
|
||||
実行後にワークスペースを検査したい場合や、既に開始済みのサンドボックスセッション上でストリーミングしたい場合に使用します。[examples/sandbox/docs/coding_task.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/docs/coding_task.py) と [examples/sandbox/docker/docker_runner.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/docker/docker_runner.py) を参照してください。
|
||||
|
||||
### セッション状態からの再開
|
||||
|
||||
`RunState` の外部で既にサンドボックス状態をシリアライズしている場合は、Runner にその状態から再接続させます。
|
||||
|
||||
```python
|
||||
from agents.run import RunConfig
|
||||
from agents.sandbox import SandboxRunConfig
|
||||
|
||||
serialized = load_saved_payload()
|
||||
restored_state = client.deserialize_session_state(serialized)
|
||||
|
||||
run_config = RunConfig(
|
||||
sandbox=SandboxRunConfig(
|
||||
client=client,
|
||||
session_state=restored_state,
|
||||
),
|
||||
)
|
||||
```
|
||||
|
||||
サンドボックス状態が独自のストレージまたはジョブシステムに存在し、`Runner` にそこから直接再開させたい場合に使用します。シリアライズ/デシリアライズのフローについては、[examples/sandbox/extensions/blaxel_runner.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/extensions/blaxel_runner.py) を参照してください。
|
||||
|
||||
### スナップショットからの開始
|
||||
|
||||
保存済みファイルと成果物から新規サンドボックスを初期化します。
|
||||
|
||||
```python
|
||||
from pathlib import Path
|
||||
|
||||
from agents.run import RunConfig
|
||||
from agents.sandbox import LocalSnapshotSpec, SandboxRunConfig
|
||||
from agents.sandbox.sandboxes.unix_local import UnixLocalSandboxClient
|
||||
|
||||
run_config = RunConfig(
|
||||
sandbox=SandboxRunConfig(
|
||||
client=UnixLocalSandboxClient(),
|
||||
snapshot=LocalSnapshotSpec(base_path=Path("/tmp/my-sandbox-snapshot")),
|
||||
),
|
||||
)
|
||||
```
|
||||
|
||||
新規実行を `agent.default_manifest` だけではなく、保存済みワークスペース内容から開始すべき場合に使用します。ローカルスナップショットフローについては [examples/sandbox/memory.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/memory.py) を、リモートスナップショットクライアントについては [examples/sandbox/sandbox_agent_with_remote_snapshot.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/sandbox_agent_with_remote_snapshot.py) を参照してください。
|
||||
|
||||
### Git からのスキル読み込み
|
||||
|
||||
ローカルスキルソースを、リポジトリ backed のものに差し替えます。
|
||||
|
||||
```python
|
||||
from agents.sandbox.capabilities import Capabilities, Skills
|
||||
from agents.sandbox.entries import GitRepo
|
||||
|
||||
capabilities = Capabilities.default() + [
|
||||
Skills(from_=GitRepo(repo="sdcoffey/tax-prep-skills", ref="main")),
|
||||
]
|
||||
```
|
||||
|
||||
スキルバンドルに独自のリリースサイクルがある場合や、サンドボックス間で共有すべき場合に使用します。[examples/sandbox/tax_prep.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/tax_prep.py) を参照してください。
|
||||
|
||||
### ツールとしての公開
|
||||
|
||||
ツールエージェントは、独自のサンドボックス境界を持つことも、親実行のライブサンドボックスを再利用することもできます。再利用は、高速な読み取り専用の探索エージェントに有用です。別のサンドボックスの作成、ハイドレーション、スナップショット作成にコストをかけずに、親が使用している正確なワークスペースを検査できます。
|
||||
|
||||
```python
|
||||
from agents import Runner
|
||||
from agents.run import RunConfig
|
||||
from agents.sandbox import FileMode, Manifest, Permissions, SandboxAgent, SandboxRunConfig, User
|
||||
from agents.sandbox.entries import Dir, File
|
||||
from agents.sandbox.sandboxes.unix_local import UnixLocalSandboxClient
|
||||
|
||||
coordinator = User(name="coordinator")
|
||||
explorer = User(name="explorer")
|
||||
|
||||
manifest = Manifest(
|
||||
users=[coordinator, explorer],
|
||||
entries={
|
||||
"pricing_packet": Dir(
|
||||
group=coordinator,
|
||||
permissions=Permissions(
|
||||
owner=FileMode.ALL,
|
||||
group=FileMode.ALL,
|
||||
other=FileMode.READ | FileMode.EXEC,
|
||||
directory=True,
|
||||
),
|
||||
children={
|
||||
"pricing.md": File(
|
||||
content=b"Pricing packet contents...",
|
||||
group=coordinator,
|
||||
permissions=Permissions(
|
||||
owner=FileMode.ALL,
|
||||
group=FileMode.ALL,
|
||||
other=FileMode.READ,
|
||||
),
|
||||
),
|
||||
},
|
||||
),
|
||||
"work": Dir(
|
||||
group=coordinator,
|
||||
permissions=Permissions(
|
||||
owner=FileMode.ALL,
|
||||
group=FileMode.ALL,
|
||||
other=FileMode.NONE,
|
||||
directory=True,
|
||||
),
|
||||
),
|
||||
},
|
||||
)
|
||||
|
||||
pricing_explorer = SandboxAgent(
|
||||
name="Pricing Explorer",
|
||||
instructions="Read `pricing_packet/` and summarize commercial risk. Do not edit files.",
|
||||
run_as=explorer,
|
||||
)
|
||||
|
||||
client = UnixLocalSandboxClient()
|
||||
sandbox = await client.create(manifest=manifest)
|
||||
|
||||
async with sandbox:
|
||||
shared_run_config = RunConfig(
|
||||
sandbox=SandboxRunConfig(session=sandbox),
|
||||
)
|
||||
|
||||
orchestrator = SandboxAgent(
|
||||
name="Revenue Operations Coordinator",
|
||||
instructions="Coordinate the review and write final notes to `work/`.",
|
||||
run_as=coordinator,
|
||||
tools=[
|
||||
pricing_explorer.as_tool(
|
||||
tool_name="review_pricing_packet",
|
||||
tool_description="Inspect the pricing packet and summarize commercial risk.",
|
||||
run_config=shared_run_config,
|
||||
max_turns=2,
|
||||
),
|
||||
],
|
||||
)
|
||||
|
||||
result = await Runner.run(
|
||||
orchestrator,
|
||||
"Review the pricing packet, then write final notes to `work/summary.md`.",
|
||||
run_config=shared_run_config,
|
||||
)
|
||||
```
|
||||
|
||||
ここでは、親エージェントは `coordinator` として実行され、探索ツールエージェントは同じライブサンドボックスセッション内で `explorer` として実行されます。`pricing_packet/` エントリは `other` ユーザーが読み取り可能であるため、explorer はすばやく検査できますが、書き込みビットは持ちません。`work/` ディレクトリは coordinator のユーザー/グループだけが利用できるため、親は最終成果物を書き込めますが、explorer は読み取り専用のままです。
|
||||
|
||||
ツールエージェントに実際の隔離が必要な場合は、代わりに独自のサンドボックス `RunConfig` を与えます。
|
||||
|
||||
```python
|
||||
from docker import from_env as docker_from_env
|
||||
|
||||
from agents.run import RunConfig
|
||||
from agents.sandbox import SandboxRunConfig
|
||||
from agents.sandbox.sandboxes.docker import DockerSandboxClient, DockerSandboxClientOptions
|
||||
|
||||
rollout_agent.as_tool(
|
||||
tool_name="review_rollout_risk",
|
||||
tool_description="Inspect the rollout packet and summarize implementation risk.",
|
||||
run_config=RunConfig(
|
||||
sandbox=SandboxRunConfig(
|
||||
client=DockerSandboxClient(docker_from_env()),
|
||||
options=DockerSandboxClientOptions(image="python:3.14-slim"),
|
||||
),
|
||||
),
|
||||
)
|
||||
```
|
||||
|
||||
ツールエージェントが自由に変更すべき場合、信頼できないコマンドを実行すべき場合、または異なるバックエンド/イメージを使用すべき場合は、別のサンドボックスを使用してください。[examples/sandbox/sandbox_agents_as_tools.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/sandbox_agents_as_tools.py) を参照してください。
|
||||
|
||||
### ローカルツールおよび MCP との組み合わせ
|
||||
|
||||
同じエージェントで通常のツールを使用しながら、サンドボックスワークスペースも維持します。
|
||||
|
||||
```python
|
||||
from agents.sandbox import SandboxAgent
|
||||
from agents.sandbox.capabilities import Shell
|
||||
|
||||
agent = SandboxAgent(
|
||||
name="Workspace reviewer",
|
||||
instructions="Inspect the workspace and call host tools when needed.",
|
||||
tools=[get_discount_approval_path],
|
||||
mcp_servers=[server],
|
||||
capabilities=[Shell()],
|
||||
)
|
||||
```
|
||||
|
||||
ワークスペース検査がエージェントの仕事の一部にすぎない場合に使用します。[examples/sandbox/sandbox_agent_with_tools.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/sandbox_agent_with_tools.py) を参照してください。
|
||||
|
||||
## メモリ
|
||||
|
||||
将来のサンドボックスエージェント実行が以前の実行から学習すべき場合は、`Memory` 機能を使用します。メモリは SDK の会話型 `Session` メモリとは別です。学びをサンドボックスワークスペース内のファイルに蒸留し、後続の実行がそれらのファイルを読み取れるようにします。
|
||||
|
||||
セットアップ、読み取り/生成の挙動、マルチターン会話、レイアウト隔離については、[エージェントメモリ](memory.md) を参照してください。
|
||||
|
||||
## 構成パターン
|
||||
|
||||
単一エージェントのパターンが明確になったら、次の設計上の問いは、より大きなシステムのどこにサンドボックス境界を置くかです。
|
||||
|
||||
サンドボックスエージェントは、引き続き SDK の他の部分と組み合わせられます。
|
||||
|
||||
- [ハンドオフ](../handoffs.md): ドキュメントの多い作業を、非サンドボックスの受付エージェントからサンドボックスレビュアーにハンドオフします。
|
||||
- [Agents as tools](../tools.md#agents-as-tools): 複数のサンドボックスエージェントをツールとして公開します。通常は各 `Agent.as_tool(...)` 呼び出しに `run_config=RunConfig(sandbox=SandboxRunConfig(...))` を渡し、各ツールが独自のサンドボックス境界を持つようにします。
|
||||
- [MCP](../mcp.md) と通常の関数ツール: サンドボックス機能は、`mcp_servers` や通常の Python ツールと共存できます。
|
||||
- [エージェントの実行](../running_agents.md): サンドボックス実行も通常の `Runner` API を使用します。
|
||||
|
||||
特に一般的なパターンは 2 つあります。
|
||||
|
||||
- ワークフローのうちワークスペース隔離が必要な部分だけを、非サンドボックスエージェントからサンドボックスエージェントにハンドオフする
|
||||
- オーケストレーターが複数のサンドボックスエージェントをツールとして公開する。通常は各 `Agent.as_tool(...)` 呼び出しごとに別々のサンドボックス `RunConfig` を使い、各ツールが独自の隔離ワークスペースを持つようにする
|
||||
|
||||
### ターンとサンドボックス実行
|
||||
|
||||
ハンドオフと agent-as-tool 呼び出しは分けて説明すると理解しやすくなります。
|
||||
|
||||
ハンドオフでは、引き続き 1 つのトップレベル実行と 1 つのトップレベルターンループがあります。アクティブなエージェントは変わりますが、実行がネストされるわけではありません。非サンドボックスの受付エージェントがサンドボックスレビュアーにハンドオフすると、同じ実行内の次のモデル呼び出しはサンドボックスエージェント用に準備され、そのサンドボックスエージェントが次のターンを担当します。言い換えると、ハンドオフは同じ実行の次のターンをどのエージェントが所有するかを変更します。[examples/sandbox/handoffs.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/handoffs.py) を参照してください。
|
||||
|
||||
`Agent.as_tool(...)` では、関係が異なります。外側のオーケストレーターは、ツールを呼び出すと決定するために外側の 1 ターンを使用し、そのツール呼び出しがサンドボックスエージェントのネストされた実行を開始します。ネストされた実行は、独自のターンループ、`max_turns`、承認、そして通常は独自のサンドボックス `RunConfig` を持ちます。1 つのネストされたターンで完了する場合もあれば、複数かかる場合もあります。外側のオーケストレーターの視点では、その作業全体が 1 つのツール呼び出しの背後にあるため、ネストされたターンは外側の実行のターンカウンターを増やしません。[examples/sandbox/sandbox_agents_as_tools.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/sandbox_agents_as_tools.py) を参照してください。
|
||||
|
||||
承認の挙動も同じ分担に従います。
|
||||
|
||||
- ハンドオフでは、サンドボックスエージェントがその実行内のアクティブエージェントになるため、承認は同じトップレベル実行に留まります
|
||||
- `Agent.as_tool(...)` では、サンドボックスツールエージェント内で発生した承認も外側の実行に表示されますが、それらは保存されたネスト実行状態から来ており、外側の実行が再開されるとネストされたサンドボックス実行を再開します
|
||||
|
||||
## 関連情報
|
||||
|
||||
- [クイックスタート](quickstart.md): 1 つのサンドボックスエージェントを実行します。
|
||||
- [サンドボックスクライアント](clients.md): ローカル、Docker、ホスト型、マウントのオプションを選択します。
|
||||
- [エージェントメモリ](memory.md): 以前のサンドボックス実行からの学びを保持し、再利用します。
|
||||
- [examples/sandbox/](https://github.com/openai/openai-agents-python/tree/main/examples/sandbox): 実行可能なローカル、コーディング、メモリ、ハンドオフ、エージェント構成のパターン。
|
||||
@@ -0,0 +1,189 @@
|
||||
---
|
||||
search:
|
||||
exclude: true
|
||||
---
|
||||
# エージェントメモリ
|
||||
|
||||
メモリにより、今後の sandbox エージェントの実行は過去の実行から学習できます。これは、メッセージ履歴を保存する SDK の会話用 [`Session`](../sessions/index.md) メモリとは別のものです。メモリは、過去の実行から得た学びを sandbox ワークスペース内のファイルに要約します。
|
||||
|
||||
!!! warning "ベータ機能"
|
||||
|
||||
Sandbox エージェントはベータ版です。一般提供までに API の詳細、デフォルト値、サポートされる機能が変更される可能性があり、時間とともにさらに高度な機能が追加されることも想定してください。
|
||||
|
||||
メモリは、今後の実行における 3 種類のコストを削減できます。
|
||||
|
||||
1. エージェントのコスト: エージェントがワークフローの完了に長い時間を要した場合、次回の実行では探索が少なくて済むはずです。これにより、トークン使用量と完了までの時間を削減できます。
|
||||
2. ユーザーのコスト: ユーザーがエージェントを修正したり好みを表明したりした場合、今後の実行でそのフィードバックを記憶できます。これにより、人による介入を削減できます。
|
||||
3. コンテキストのコスト: エージェントが以前にタスクを完了していて、ユーザーがそのタスクを発展させたい場合、ユーザーは以前のスレッドを探したり、すべてのコンテキストを再入力したりする必要がないはずです。これにより、タスク説明を短くできます。
|
||||
|
||||
バグを修正し、メモリを生成し、スナップショットを再開し、そのメモリを後続の検証実行で使用する、2 回の実行からなる完全なコード例については、[examples/sandbox/memory.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/memory.py) を参照してください。独立したメモリレイアウトを持つマルチターン、マルチエージェントのコード例については、[examples/sandbox/memory_multi_agent_multiturn.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/memory_multi_agent_multiturn.py) を参照してください。
|
||||
|
||||
## メモリの有効化
|
||||
|
||||
sandbox エージェントに機能として `Memory()` を追加します。
|
||||
|
||||
```python
|
||||
from pathlib import Path
|
||||
import tempfile
|
||||
|
||||
from agents.sandbox import LocalSnapshotSpec, SandboxAgent
|
||||
from agents.sandbox.capabilities import Filesystem, Memory, Shell
|
||||
|
||||
agent = SandboxAgent(
|
||||
name="Memory-enabled reviewer",
|
||||
instructions="Inspect the workspace and preserve useful lessons for follow-up runs.",
|
||||
capabilities=[Memory(), Filesystem(), Shell()],
|
||||
)
|
||||
|
||||
with tempfile.TemporaryDirectory(prefix="sandbox-memory-example-") as snapshot_dir:
|
||||
sandbox = await client.create(
|
||||
manifest=manifest,
|
||||
snapshot=LocalSnapshotSpec(base_path=Path(snapshot_dir)),
|
||||
)
|
||||
```
|
||||
|
||||
読み取りが有効な場合、`Memory()` には `Shell()` が必要です。これにより、挿入されたサマリーだけでは不十分なときに、エージェントがメモリファイルを読み取り、検索できます。ライブメモリ更新が有効な場合(デフォルト)、`Filesystem()` も必要です。これにより、エージェントが古くなったメモリを発見した場合や、ユーザーがメモリの更新を依頼した場合に、`memories/MEMORY.md` を更新できます。
|
||||
|
||||
デフォルトでは、メモリアーティファクトは sandbox ワークスペースの `memories/` 配下に保存されます。後の実行で再利用するには、同じライブ sandbox セッションを維持するか、永続化されたセッション状態またはスナップショットから再開することで、設定済みのメモリディレクトリ全体を保持して再利用してください。新しい空の sandbox は空のメモリで開始されます。
|
||||
|
||||
`Memory()` は、メモリの読み取りと生成の両方を有効にします。メモリを読み取るが新しいメモリは生成すべきでないエージェントには、`Memory(generate=None)` を使用します。たとえば、内部エージェント、サブエージェント、チェッカー、または実行から得られるシグナルが多くない 1 回限りのツールエージェントです。後で使うメモリを生成する必要はあるものの、ユーザーが既存メモリによる影響を望まない場合は、`Memory(read=None)` を使用します。
|
||||
|
||||
## メモリの読み取り
|
||||
|
||||
メモリ読み取りでは段階的開示を使用します。実行の開始時に、SDK は一般的に役立つヒント、ユーザーの好み、利用可能なメモリの小さなサマリー(`memory_summary.md`)を、エージェントの developer プロンプトに挿入します。これにより、エージェントは過去の作業が関連しそうかどうかを判断するのに十分なコンテキストを得られます。
|
||||
|
||||
過去の作業が関連しそうな場合、エージェントは現在のタスクからキーワードを抽出して、設定されたメモリインデックス(`memories_dir` 配下の `MEMORY.md`)を検索します。より詳細が必要な場合にのみ、設定された `rollout_summaries/` ディレクトリ配下にある対応する過去のロールアウトサマリーを開きます。
|
||||
|
||||
メモリは古くなることがあります。エージェントには、メモリをガイダンスとしてのみ扱い、現在の環境を信頼するよう指示されています。デフォルトでは、メモリ読み取りでは `live_update` が有効です。そのため、エージェントが古くなったメモリを発見した場合、同じ実行内で設定済みの `MEMORY.md` を更新できます。実行中にメモリを読み取るが変更してほしくない場合、たとえばレイテンシに敏感な実行では、ライブ更新を無効にしてください。
|
||||
|
||||
## メモリの生成
|
||||
|
||||
実行が完了すると、sandbox ランタイムはその実行セグメントを会話ファイルに追記します。蓄積された会話ファイルは、sandbox セッションが閉じられるときに処理されます。
|
||||
|
||||
メモリ生成には 2 つのフェーズがあります。
|
||||
|
||||
1. フェーズ 1: 会話の抽出。メモリ生成モデルが、蓄積された 1 つの会話ファイルを処理し、会話サマリーを生成します。system、developer、reasoning のコンテンツは省略されます。会話が長すぎる場合は、先頭と末尾を保持したうえで、コンテキストウィンドウに収まるよう切り詰められます。また、未加工のメモリ抽出も生成します。これは、フェーズ 2 が統合できる会話からの簡潔なメモです。
|
||||
2. フェーズ 2: レイアウトの統合。統合エージェントは、1 つのメモリレイアウトに対応する未加工のメモリを読み取り、より多くの根拠が必要な場合は会話サマリーを開き、パターンを `MEMORY.md` と `memory_summary.md` に抽出します。
|
||||
|
||||
デフォルトのワークスペースレイアウトは次のとおりです。
|
||||
|
||||
```text
|
||||
workspace/
|
||||
├── sessions/
|
||||
│ └── <rollout-id>.jsonl
|
||||
└── memories/
|
||||
├── memory_summary.md
|
||||
├── MEMORY.md
|
||||
├── raw_memories.md (intermediate)
|
||||
├── phase_two_selection.json (intermediate)
|
||||
├── raw_memories/ (intermediate)
|
||||
│ └── <rollout-id>.md
|
||||
├── rollout_summaries/
|
||||
│ └── <rollout-id>_<slug>.md
|
||||
└── skills/
|
||||
```
|
||||
|
||||
`MemoryGenerateConfig` でメモリ生成を設定できます。
|
||||
|
||||
```python
|
||||
from agents.sandbox import MemoryGenerateConfig
|
||||
from agents.sandbox.capabilities import Memory
|
||||
|
||||
memory = Memory(
|
||||
generate=MemoryGenerateConfig(
|
||||
max_raw_memories_for_consolidation=128,
|
||||
extra_prompt="Pay extra attention to what made the customer more satisfied or annoyed",
|
||||
),
|
||||
)
|
||||
```
|
||||
|
||||
`extra_prompt` を使用して、GTM エージェント向けの顧客や会社の詳細など、ユースケースで最も重要なシグナルをメモリ生成器に伝えます。
|
||||
|
||||
最近の未加工メモリが `max_raw_memories_for_consolidation`(デフォルトは 256)を超える場合、フェーズ 2 は最新の会話のメモリだけを保持し、古いものを削除します。新しさは、会話が最後に更新された時刻に基づきます。この忘却メカニズムにより、メモリが最新の環境を反映しやすくなります。
|
||||
|
||||
## マルチターン会話
|
||||
|
||||
マルチターンの sandbox チャットでは、同じライブ sandbox セッションとともに通常の SDK `Session` を使用します。
|
||||
|
||||
```python
|
||||
from agents import Runner, SQLiteSession
|
||||
from agents.run import RunConfig
|
||||
from agents.sandbox import SandboxRunConfig
|
||||
|
||||
conversation_session = SQLiteSession("gtm-q2-pipeline-review")
|
||||
sandbox = await client.create(manifest=agent.default_manifest)
|
||||
|
||||
async with sandbox:
|
||||
run_config = RunConfig(
|
||||
sandbox=SandboxRunConfig(session=sandbox),
|
||||
workflow_name="GTM memory example",
|
||||
)
|
||||
await Runner.run(
|
||||
agent,
|
||||
"Analyze data/leads.csv and identify one promising GTM segment.",
|
||||
session=conversation_session,
|
||||
run_config=run_config,
|
||||
)
|
||||
await Runner.run(
|
||||
agent,
|
||||
"Using that analysis, write a short outreach hypothesis.",
|
||||
session=conversation_session,
|
||||
run_config=run_config,
|
||||
)
|
||||
```
|
||||
|
||||
どちらの実行も、同じ SDK 会話セッション(`session=conversation_session`)を渡すため、1 つのメモリ会話ファイルに追記され、したがって同じ `session.session_id` を共有します。これはライブワークスペースを識別する sandbox(`sandbox`)とは異なります。`sandbox` はメモリ会話 ID としては使用されません。sandbox セッションが閉じられると、フェーズ 1 は蓄積された会話を参照するため、2 つの孤立したターンではなく、やり取り全体からメモリを抽出できます。
|
||||
|
||||
複数の `Runner.run(...)` 呼び出しを 1 つのメモリ会話にしたい場合は、それらの呼び出し全体で安定した識別子を渡してください。メモリが実行を会話に関連付けるときは、次の順序で解決します。
|
||||
|
||||
1. `conversation_id`(`Runner.run(...)` に渡した場合)
|
||||
2. `session.session_id`(`SQLiteSession` などの SDK `Session` を渡した場合)
|
||||
3. `RunConfig.group_id`(上記のどちらも存在しない場合)
|
||||
4. 生成された実行ごとの ID(安定した識別子が存在しない場合)
|
||||
|
||||
## エージェントごとのメモリ分離における異なるレイアウトの利用
|
||||
|
||||
メモリの分離はエージェント名ではなく `MemoryLayoutConfig` に基づきます。同じレイアウトと同じメモリ会話 ID を持つエージェントは、1 つのメモリ会話と 1 つの統合済みメモリを共有します。異なるレイアウトを持つエージェントは、同じ sandbox ワークスペースを共有している場合でも、別々のロールアウトファイル、未加工メモリ、`MEMORY.md`、`memory_summary.md` を保持します。
|
||||
|
||||
複数のエージェントが 1 つの sandbox を共有するものの、メモリは共有すべきでない場合は、別々のレイアウトを使用します。
|
||||
|
||||
```python
|
||||
from agents import SQLiteSession
|
||||
from agents.sandbox import MemoryLayoutConfig, SandboxAgent
|
||||
from agents.sandbox.capabilities import Filesystem, Memory, Shell
|
||||
|
||||
gtm_agent = SandboxAgent(
|
||||
name="GTM reviewer",
|
||||
instructions="Analyze GTM workspace data and write concise recommendations.",
|
||||
capabilities=[
|
||||
Memory(
|
||||
layout=MemoryLayoutConfig(
|
||||
memories_dir="memories/gtm",
|
||||
sessions_dir="sessions/gtm",
|
||||
)
|
||||
),
|
||||
Filesystem(),
|
||||
Shell(),
|
||||
],
|
||||
)
|
||||
|
||||
engineering_agent = SandboxAgent(
|
||||
name="Engineering reviewer",
|
||||
instructions="Inspect engineering workspaces and summarize fixes and risks.",
|
||||
capabilities=[
|
||||
Memory(
|
||||
layout=MemoryLayoutConfig(
|
||||
memories_dir="memories/engineering",
|
||||
sessions_dir="sessions/engineering",
|
||||
)
|
||||
),
|
||||
Filesystem(),
|
||||
Shell(),
|
||||
],
|
||||
)
|
||||
|
||||
gtm_session = SQLiteSession("gtm-q2-pipeline-review")
|
||||
engineering_session = SQLiteSession("eng-invoice-test-fix")
|
||||
```
|
||||
|
||||
これにより、GTM 分析がエンジニアリングのバグ修正メモリに統合されたり、その逆が起きたりすることを防げます。
|
||||
@@ -0,0 +1,117 @@
|
||||
---
|
||||
search:
|
||||
exclude: true
|
||||
---
|
||||
# クイックスタート
|
||||
|
||||
!!! warning "ベータ機能"
|
||||
|
||||
サンドボックスエージェントはベータ版です。一般提供前に API の詳細、デフォルト値、サポートされる機能が変更される可能性があり、今後より高度な機能が追加されることが想定されます。
|
||||
|
||||
最新のエージェントは、ファイルシステム上の実際のファイルを操作できるときに最も効果を発揮します。Agents SDK の **サンドボックスエージェント** は、モデルに永続的なワークスペースを提供し、大規模なドキュメントセットの検索、ファイル編集、コマンド実行、成果物の生成、保存済みのサンドボックス状態からの作業再開を可能にします。
|
||||
|
||||
SDK は、ファイルのステージング、ファイルシステムツール、シェルアクセス、サンドボックスのライフサイクル、スナップショット、プロバイダー固有の連携を自分で組み合わせる必要なく、その実行ハーネスを提供します。通常の `Agent` と `Runner` のフローはそのままに、ワークスペース用の `Manifest`、サンドボックスネイティブツール用の機能、作業の実行場所を指定する `SandboxRunConfig` を追加します。
|
||||
|
||||
## 前提条件
|
||||
|
||||
- Python 3.10 以上
|
||||
- OpenAI Agents SDK に関する基本的な理解
|
||||
- サンドボックスクライアント。ローカル開発では、`UnixLocalSandboxClient` から始めてください。
|
||||
|
||||
## インストール
|
||||
|
||||
SDK をまだインストールしていない場合:
|
||||
|
||||
```bash
|
||||
pip install openai-agents
|
||||
```
|
||||
|
||||
Docker ベースのサンドボックスの場合:
|
||||
|
||||
```bash
|
||||
pip install "openai-agents[docker]"
|
||||
```
|
||||
|
||||
## ローカルサンドボックスエージェントの作成
|
||||
|
||||
この例では、ローカルリポジトリを `repo/` 配下にステージングし、ローカルスキルを遅延ロードし、ランナーが実行用の Unix ローカルサンドボックスセッションを作成できるようにします。
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
from pathlib import Path
|
||||
|
||||
from agents import Runner
|
||||
from agents.run import RunConfig
|
||||
from agents.sandbox import Manifest, SandboxAgent, SandboxRunConfig
|
||||
from agents.sandbox.capabilities import Capabilities, LocalDirLazySkillSource, Skills
|
||||
from agents.sandbox.entries import LocalDir
|
||||
from agents.sandbox.sandboxes.unix_local import UnixLocalSandboxClient
|
||||
|
||||
EXAMPLE_DIR = Path(__file__).resolve().parent
|
||||
HOST_REPO_DIR = EXAMPLE_DIR / "repo"
|
||||
HOST_SKILLS_DIR = EXAMPLE_DIR / "skills"
|
||||
|
||||
|
||||
def build_agent(model: str) -> SandboxAgent[None]:
|
||||
return SandboxAgent(
|
||||
name="Sandbox engineer",
|
||||
model=model,
|
||||
instructions=(
|
||||
"Read `repo/task.md` before editing files. Stay grounded in the repository, preserve "
|
||||
"existing behavior, and mention the exact verification command you ran. "
|
||||
"If you edit files with apply_patch, paths are relative to the sandbox workspace root."
|
||||
),
|
||||
default_manifest=Manifest(
|
||||
entries={
|
||||
"repo": LocalDir(src=HOST_REPO_DIR),
|
||||
}
|
||||
),
|
||||
capabilities=Capabilities.default() + [
|
||||
Skills(
|
||||
lazy_from=LocalDirLazySkillSource(
|
||||
# This is a host path read by the SDK process.
|
||||
# Requested skills are copied into `skills_path` in the sandbox.
|
||||
source=LocalDir(src=HOST_SKILLS_DIR),
|
||||
)
|
||||
),
|
||||
],
|
||||
)
|
||||
|
||||
|
||||
async def main() -> None:
|
||||
result = await Runner.run(
|
||||
build_agent("gpt-5.5"),
|
||||
"Open `repo/task.md`, fix the issue, run the targeted test, and summarize the change.",
|
||||
run_config=RunConfig(
|
||||
sandbox=SandboxRunConfig(client=UnixLocalSandboxClient()),
|
||||
workflow_name="Sandbox coding example",
|
||||
),
|
||||
)
|
||||
print(result.final_output)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
[examples/sandbox/docs/coding_task.py](https://github.com/openai/openai-agents-python/blob/main/examples/sandbox/docs/coding_task.py) を参照してください。この例は小さなシェルベースのリポジトリを使用しているため、Unix ローカル実行間で決定論的に検証できます。
|
||||
|
||||
## 主な選択肢
|
||||
|
||||
基本的な実行が動作したら、次に多くの人が検討する選択肢は次のとおりです。
|
||||
|
||||
- `default_manifest`: 新しいサンドボックスセッション用のファイル、リポジトリ、ディレクトリ、マウント
|
||||
- `instructions`: 複数のプロンプトにわたって適用すべき短いワークフロールール
|
||||
- `base_instructions`: SDK のサンドボックスプロンプトを置き換えるための高度なエスケープハッチ
|
||||
- `capabilities`: ファイルシステム編集 / 画像検査、シェル、スキル、メモリ、コンパクションなどのサンドボックスネイティブツール
|
||||
- `run_as`: モデル向けツールで使用するサンドボックスのユーザー ID
|
||||
- `SandboxRunConfig.client`: サンドボックスバックエンド
|
||||
- `SandboxRunConfig.session`、`session_state`、または `snapshot`: 後続の実行が以前の作業に再接続する方法
|
||||
|
||||
## 次のステップ
|
||||
|
||||
- [概念](sandbox/guide.md): マニフェスト、機能、権限、スナップショット、実行設定、構成パターンを理解します。
|
||||
- [サンドボックスクライアント](sandbox/clients.md): Unix ローカル、Docker、ホスト型プロバイダー、マウント戦略を選択します。
|
||||
- [エージェントメモリ](sandbox/memory.md): 以前のサンドボックス実行から得た知見を保存し、再利用します。
|
||||
|
||||
シェルアクセスがたまに使うツールの 1 つにすぎない場合は、[ツールガイド](tools.md) のホスト型シェルから始めてください。ワークスペース分離、サンドボックスクライアントの選択、またはサンドボックスセッションの再開動作が設計に含まれる場合は、サンドボックスエージェントを選択してください。
|
||||
@@ -0,0 +1,459 @@
|
||||
---
|
||||
search:
|
||||
exclude: true
|
||||
---
|
||||
# セッション
|
||||
|
||||
Agents SDK は、複数のエージェント実行にわたって会話履歴を自動で維持する組み込みのセッションメモリを提供し、ターン間で手動で `.to_input_list()` を扱う必要をなくします。
|
||||
|
||||
セッションは特定のセッションに対する会話履歴を保存し、明示的な手動メモリ管理なしでエージェントがコンテキストを維持できるようにします。これは、エージェントに過去のやり取りを記憶させたいチャットアプリケーションやマルチターンの会話を構築する際に特に有用です。
|
||||
|
||||
## クイックスタート
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner, SQLiteSession
|
||||
|
||||
# Create agent
|
||||
agent = Agent(
|
||||
name="Assistant",
|
||||
instructions="Reply very concisely.",
|
||||
)
|
||||
|
||||
# Create a session instance with a session ID
|
||||
session = SQLiteSession("conversation_123")
|
||||
|
||||
# First turn
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
"What city is the Golden Gate Bridge in?",
|
||||
session=session
|
||||
)
|
||||
print(result.final_output) # "San Francisco"
|
||||
|
||||
# Second turn - agent automatically remembers previous context
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
"What state is it in?",
|
||||
session=session
|
||||
)
|
||||
print(result.final_output) # "California"
|
||||
|
||||
# Also works with synchronous runner
|
||||
result = Runner.run_sync(
|
||||
agent,
|
||||
"What's the population?",
|
||||
session=session
|
||||
)
|
||||
print(result.final_output) # "Approximately 39 million"
|
||||
```
|
||||
|
||||
## 仕組み
|
||||
|
||||
セッションメモリが有効な場合:
|
||||
|
||||
1. **各実行の前**: ランナーはセッションの会話履歴を自動的に取得し、入力アイテムの前に付加します。
|
||||
2. **各実行の後**: 実行中に生成されたすべての新しいアイテム (ユーザー入力、アシスタントの応答、ツール呼び出しなど) は自動的にセッションに保存されます。
|
||||
3. **コンテキスト保持**: 同一セッションでの後続の実行には完全な会話履歴が含まれ、エージェントはコンテキストを維持できます。
|
||||
|
||||
これにより、ターン間で `.to_input_list()` を手動で呼び出して会話状態を管理する必要がなくなります。
|
||||
|
||||
## メモリ操作
|
||||
|
||||
### 基本操作
|
||||
|
||||
セッションは会話履歴を管理するためにいくつかの操作をサポートします:
|
||||
|
||||
```python
|
||||
from agents import SQLiteSession
|
||||
|
||||
session = SQLiteSession("user_123", "conversations.db")
|
||||
|
||||
# Get all items in a session
|
||||
items = await session.get_items()
|
||||
|
||||
# Add new items to a session
|
||||
new_items = [
|
||||
{"role": "user", "content": "Hello"},
|
||||
{"role": "assistant", "content": "Hi there!"}
|
||||
]
|
||||
await session.add_items(new_items)
|
||||
|
||||
# Remove and return the most recent item
|
||||
last_item = await session.pop_item()
|
||||
print(last_item) # {"role": "assistant", "content": "Hi there!"}
|
||||
|
||||
# Clear all items from a session
|
||||
await session.clear_session()
|
||||
```
|
||||
|
||||
### 修正のための pop_item の使用
|
||||
|
||||
会話内の最後のアイテムを取り消したり修正したい場合、`pop_item` メソッドが特に便利です:
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner, SQLiteSession
|
||||
|
||||
agent = Agent(name="Assistant")
|
||||
session = SQLiteSession("correction_example")
|
||||
|
||||
# Initial conversation
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
"What's 2 + 2?",
|
||||
session=session
|
||||
)
|
||||
print(f"Agent: {result.final_output}")
|
||||
|
||||
# User wants to correct their question
|
||||
assistant_item = await session.pop_item() # Remove agent's response
|
||||
user_item = await session.pop_item() # Remove user's question
|
||||
|
||||
# Ask a corrected question
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
"What's 2 + 3?",
|
||||
session=session
|
||||
)
|
||||
print(f"Agent: {result.final_output}")
|
||||
```
|
||||
|
||||
## メモリオプション
|
||||
|
||||
### メモリなし (デフォルト)
|
||||
|
||||
```python
|
||||
# Default behavior - no session memory
|
||||
result = await Runner.run(agent, "Hello")
|
||||
```
|
||||
|
||||
### OpenAI Conversations API メモリ
|
||||
|
||||
自前のデータベースを管理せずに [会話状態](https://platform.openai.com/docs/guides/conversation-state?api-mode=responses#using-the-conversations-api) を永続化するには、[OpenAI Conversations API](https://platform.openai.com/docs/api-reference/conversations/create) を使用します。これは、会話履歴の保存に OpenAI がホストするインフラストラクチャに既に依存している場合に役立ちます。
|
||||
|
||||
```python
|
||||
from agents import OpenAIConversationsSession
|
||||
|
||||
session = OpenAIConversationsSession()
|
||||
|
||||
# Optionally resume a previous conversation by passing a conversation ID
|
||||
# session = OpenAIConversationsSession(conversation_id="conv_123")
|
||||
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
"Hello",
|
||||
session=session,
|
||||
)
|
||||
```
|
||||
|
||||
### SQLite メモリ
|
||||
|
||||
```python
|
||||
from agents import SQLiteSession
|
||||
|
||||
# In-memory database (lost when process ends)
|
||||
session = SQLiteSession("user_123")
|
||||
|
||||
# Persistent file-based database
|
||||
session = SQLiteSession("user_123", "conversations.db")
|
||||
|
||||
# Use the session
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
"Hello",
|
||||
session=session
|
||||
)
|
||||
```
|
||||
|
||||
### 複数セッション
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner, SQLiteSession
|
||||
|
||||
agent = Agent(name="Assistant")
|
||||
|
||||
# Different sessions maintain separate conversation histories
|
||||
session_1 = SQLiteSession("user_123", "conversations.db")
|
||||
session_2 = SQLiteSession("user_456", "conversations.db")
|
||||
|
||||
result1 = await Runner.run(
|
||||
agent,
|
||||
"Hello",
|
||||
session=session_1
|
||||
)
|
||||
result2 = await Runner.run(
|
||||
agent,
|
||||
"Hello",
|
||||
session=session_2
|
||||
)
|
||||
```
|
||||
|
||||
### SQLAlchemy ベースのセッション
|
||||
|
||||
より高度なユースケースでは、SQLAlchemy ベースのセッションバックエンドを使用できます。これにより、セッションストレージに SQLAlchemy がサポートする任意のデータベース (PostgreSQL、MySQL、SQLite など) を使用できます。
|
||||
|
||||
**例 1: `from_url` を使ったインメモリ SQLite**
|
||||
|
||||
これは最も簡単な開始方法で、開発やテストに最適です。
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
from agents import Agent, Runner
|
||||
from agents.extensions.memory.sqlalchemy_session import SQLAlchemySession
|
||||
|
||||
async def main():
|
||||
agent = Agent("Assistant")
|
||||
session = SQLAlchemySession.from_url(
|
||||
"user-123",
|
||||
url="sqlite+aiosqlite:///:memory:",
|
||||
create_tables=True, # Auto-create tables for the demo
|
||||
)
|
||||
|
||||
result = await Runner.run(agent, "Hello", session=session)
|
||||
|
||||
if __name__ == "__main__":
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
**例 2: 既存の SQLAlchemy エンジンを使用**
|
||||
|
||||
本番アプリケーションでは、すでに SQLAlchemy の `AsyncEngine` インスタンスを持っている可能性が高いです。これをそのままセッションに渡せます。
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
from agents import Agent, Runner
|
||||
from agents.extensions.memory.sqlalchemy_session import SQLAlchemySession
|
||||
from sqlalchemy.ext.asyncio import create_async_engine
|
||||
|
||||
async def main():
|
||||
# In your application, you would use your existing engine
|
||||
engine = create_async_engine("sqlite+aiosqlite:///conversations.db")
|
||||
|
||||
agent = Agent("Assistant")
|
||||
session = SQLAlchemySession(
|
||||
"user-456",
|
||||
engine=engine,
|
||||
create_tables=True, # Auto-create tables for the demo
|
||||
)
|
||||
|
||||
result = await Runner.run(agent, "Hello", session=session)
|
||||
print(result.final_output)
|
||||
|
||||
await engine.dispose()
|
||||
|
||||
if __name__ == "__main__":
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
### 暗号化セッション
|
||||
|
||||
保存時に会話データの暗号化が必要なアプリケーションでは、`EncryptedSession` を使用して任意のセッションバックエンドを透過的な暗号化と自動 TTL ベースの有効期限でラップできます。これには `encrypt` エクストラが必要です: `pip install openai-agents[encrypt]`。
|
||||
|
||||
`EncryptedSession` は、セッションごとのキー導出 (HKDF) を用いた Fernet 暗号化を使用し、古いメッセージの自動期限切れをサポートします。アイテムが TTL を超えると、取得時に静かにスキップされます。
|
||||
|
||||
**例: SQLAlchemy セッションデータの暗号化**
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
from agents import Agent, Runner
|
||||
from agents.extensions.memory import EncryptedSession, SQLAlchemySession
|
||||
|
||||
async def main():
|
||||
# Create underlying session (works with any SessionABC implementation)
|
||||
underlying_session = SQLAlchemySession.from_url(
|
||||
session_id="user-123",
|
||||
url="postgresql+asyncpg://app:secret@db.example.com/agents",
|
||||
create_tables=True,
|
||||
)
|
||||
|
||||
# Wrap with encryption and TTL-based expiration
|
||||
session = EncryptedSession(
|
||||
session_id="user-123",
|
||||
underlying_session=underlying_session,
|
||||
encryption_key="your-encryption-key", # Use a secure key from your secrets management
|
||||
ttl=600, # 10 minutes - items older than this are silently skipped
|
||||
)
|
||||
|
||||
agent = Agent("Assistant")
|
||||
result = await Runner.run(agent, "Hello", session=session)
|
||||
print(result.final_output)
|
||||
|
||||
if __name__ == "__main__":
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
**主な特長:**
|
||||
|
||||
- **透過的な暗号化**: 保存前にすべてのセッションアイテムを自動的に暗号化し、取得時に復号化します
|
||||
- **セッションごとのキー導出**: セッション ID をソルトとした HKDF で一意の暗号鍵を導出します
|
||||
- **TTL ベースの有効期限**: 設定可能な有効期間に基づいて古いメッセージを自動的に期限切れにします (デフォルト: 10 分)
|
||||
- **柔軟な鍵入力**: Fernet キーまたは生の文字列のいずれも暗号鍵として受け付けます
|
||||
- **任意のセッションをラップ**: SQLite、SQLAlchemy、またはカスタムセッション実装で動作します
|
||||
|
||||
!!! warning "重要なセキュリティに関する注意"
|
||||
|
||||
- 暗号鍵は安全に保管してください (例: 環境変数、シークレットマネージャー)
|
||||
- 期限切れトークンの拒否はアプリケーション サーバーのシステムクロックに基づきます。正当なトークンがクロックずれにより拒否されないよう、すべてのサーバーが NTP で時刻同期されていることを確認してください
|
||||
- 基盤となるセッションは暗号化済みデータを保存し続けるため、データベース インフラストラクチャの管理権限は保持されます
|
||||
|
||||
|
||||
## カスタムメモリ実装
|
||||
|
||||
[`Session`][agents.memory.session.Session] プロトコルに従うクラスを作成することで、独自のセッションメモリを実装できます:
|
||||
|
||||
```python
|
||||
from agents.memory.session import SessionABC
|
||||
from agents.items import TResponseInputItem
|
||||
from typing import List
|
||||
|
||||
class MyCustomSession(SessionABC):
|
||||
"""Custom session implementation following the Session protocol."""
|
||||
|
||||
def __init__(self, session_id: str):
|
||||
self.session_id = session_id
|
||||
# Your initialization here
|
||||
|
||||
async def get_items(self, limit: int | None = None) -> List[TResponseInputItem]:
|
||||
"""Retrieve conversation history for this session."""
|
||||
# Your implementation here
|
||||
pass
|
||||
|
||||
async def add_items(self, items: List[TResponseInputItem]) -> None:
|
||||
"""Store new items for this session."""
|
||||
# Your implementation here
|
||||
pass
|
||||
|
||||
async def pop_item(self) -> TResponseInputItem | None:
|
||||
"""Remove and return the most recent item from this session."""
|
||||
# Your implementation here
|
||||
pass
|
||||
|
||||
async def clear_session(self) -> None:
|
||||
"""Clear all items for this session."""
|
||||
# Your implementation here
|
||||
pass
|
||||
|
||||
# Use your custom session
|
||||
agent = Agent(name="Assistant")
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
"Hello",
|
||||
session=MyCustomSession("my_session")
|
||||
)
|
||||
```
|
||||
|
||||
## セッション管理
|
||||
|
||||
### セッション ID の命名
|
||||
|
||||
会話の整理に役立つわかりやすいセッション ID を使用します:
|
||||
|
||||
- ユーザー基準: `"user_12345"`
|
||||
- スレッド基準: `"thread_abc123"`
|
||||
- コンテキスト基準: `"support_ticket_456"`
|
||||
|
||||
### メモリ永続化
|
||||
|
||||
- 一時的な会話にはインメモリ SQLite (`SQLiteSession("session_id")`) を使用
|
||||
- 永続的な会話にはファイルベース SQLite (`SQLiteSession("session_id", "path/to/db.sqlite")`) を使用
|
||||
- 既存のデータベースを持つ本番システムには SQLAlchemy ベースのセッション (`SQLAlchemySession("session_id", engine=engine, create_tables=True)`) を使用
|
||||
- 履歴を OpenAI Conversations API に保存したい場合は OpenAI がホストするストレージ (`OpenAIConversationsSession()`) を使用
|
||||
- 透過的な暗号化と TTL ベースの有効期限で任意のセッションをラップするには暗号化セッション (`EncryptedSession(session_id, underlying_session, encryption_key)`) を使用
|
||||
- さらに高度なユースケース向けに、他の本番システム (Redis、Django など) 用のカスタムセッションバックエンドの実装を検討
|
||||
|
||||
### セッション管理
|
||||
|
||||
```python
|
||||
# Clear a session when conversation should start fresh
|
||||
await session.clear_session()
|
||||
|
||||
# Different agents can share the same session
|
||||
support_agent = Agent(name="Support")
|
||||
billing_agent = Agent(name="Billing")
|
||||
session = SQLiteSession("user_123")
|
||||
|
||||
# Both agents will see the same conversation history
|
||||
result1 = await Runner.run(
|
||||
support_agent,
|
||||
"Help me with my account",
|
||||
session=session
|
||||
)
|
||||
result2 = await Runner.run(
|
||||
billing_agent,
|
||||
"What are my charges?",
|
||||
session=session
|
||||
)
|
||||
```
|
||||
|
||||
## 完全な例
|
||||
|
||||
セッションメモリの動作を示す完全な例です:
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
from agents import Agent, Runner, SQLiteSession
|
||||
|
||||
|
||||
async def main():
|
||||
# Create an agent
|
||||
agent = Agent(
|
||||
name="Assistant",
|
||||
instructions="Reply very concisely.",
|
||||
)
|
||||
|
||||
# Create a session instance that will persist across runs
|
||||
session = SQLiteSession("conversation_123", "conversation_history.db")
|
||||
|
||||
print("=== Sessions Example ===")
|
||||
print("The agent will remember previous messages automatically.\n")
|
||||
|
||||
# First turn
|
||||
print("First turn:")
|
||||
print("User: What city is the Golden Gate Bridge in?")
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
"What city is the Golden Gate Bridge in?",
|
||||
session=session
|
||||
)
|
||||
print(f"Assistant: {result.final_output}")
|
||||
print()
|
||||
|
||||
# Second turn - the agent will remember the previous conversation
|
||||
print("Second turn:")
|
||||
print("User: What state is it in?")
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
"What state is it in?",
|
||||
session=session
|
||||
)
|
||||
print(f"Assistant: {result.final_output}")
|
||||
print()
|
||||
|
||||
# Third turn - continuing the conversation
|
||||
print("Third turn:")
|
||||
print("User: What's the population of that state?")
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
"What's the population of that state?",
|
||||
session=session
|
||||
)
|
||||
print(f"Assistant: {result.final_output}")
|
||||
print()
|
||||
|
||||
print("=== Conversation Complete ===")
|
||||
print("Notice how the agent remembered the context from previous turns!")
|
||||
print("Sessions automatically handles conversation history.")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
## API リファレンス
|
||||
|
||||
詳細な API ドキュメントは以下をご覧ください:
|
||||
|
||||
- [`Session`][agents.memory.Session] - プロトコルインターフェース
|
||||
- [`SQLiteSession`][agents.memory.SQLiteSession] - SQLite 実装
|
||||
- [`OpenAIConversationsSession`](ref/memory/openai_conversations_session.md) - OpenAI Conversations API 実装
|
||||
- [`SQLAlchemySession`][agents.extensions.memory.sqlalchemy_session.SQLAlchemySession] - SQLAlchemy ベースの実装
|
||||
- [`EncryptedSession`][agents.extensions.memory.encrypt_session.EncryptedSession] - TTL 付き暗号化セッションラッパー
|
||||
@@ -0,0 +1,307 @@
|
||||
---
|
||||
search:
|
||||
exclude: true
|
||||
---
|
||||
# 高度な SQLite セッション
|
||||
|
||||
`AdvancedSQLiteSession` は基本的な `SQLiteSession` の拡張版であり、会話の分岐、詳細な使用状況分析、構造化された会話クエリなど、高度な会話管理機能を提供します。
|
||||
|
||||
## 機能
|
||||
|
||||
- **会話の分岐**: 任意のユーザーメッセージから別の会話パスを作成します
|
||||
- **使用状況の追跡**: ターンごとの詳細なトークン使用状況分析を、完全な JSON 内訳付きで提供します
|
||||
- **構造化クエリ**: ターン別の会話、ツール使用状況の統計などを取得します
|
||||
- **ブランチ管理**: 独立したブランチ切り替えと管理を行います
|
||||
- **メッセージ構造メタデータ**: メッセージタイプ、ツール使用、会話フローを追跡します
|
||||
|
||||
## クイックスタート
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner
|
||||
from agents.extensions.memory import AdvancedSQLiteSession
|
||||
|
||||
# Create agent
|
||||
agent = Agent(
|
||||
name="Assistant",
|
||||
instructions="Reply very concisely.",
|
||||
)
|
||||
|
||||
# Create an advanced session
|
||||
session = AdvancedSQLiteSession(
|
||||
session_id="conversation_123",
|
||||
db_path="conversations.db",
|
||||
create_tables=True
|
||||
)
|
||||
|
||||
# First conversation turn
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
"What city is the Golden Gate Bridge in?",
|
||||
session=session
|
||||
)
|
||||
print(result.final_output) # "San Francisco"
|
||||
|
||||
# IMPORTANT: Store usage data
|
||||
await session.store_run_usage(result)
|
||||
|
||||
# Continue conversation
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
"What state is it in?",
|
||||
session=session
|
||||
)
|
||||
print(result.final_output) # "California"
|
||||
await session.store_run_usage(result)
|
||||
```
|
||||
|
||||
## 初期化
|
||||
|
||||
```python
|
||||
from agents.extensions.memory import AdvancedSQLiteSession
|
||||
|
||||
# Basic initialization
|
||||
session = AdvancedSQLiteSession(
|
||||
session_id="my_conversation",
|
||||
create_tables=True # Auto-create advanced tables
|
||||
)
|
||||
|
||||
# With persistent storage
|
||||
session = AdvancedSQLiteSession(
|
||||
session_id="user_123",
|
||||
db_path="path/to/conversations.db",
|
||||
create_tables=True
|
||||
)
|
||||
|
||||
# With custom logger
|
||||
import logging
|
||||
logger = logging.getLogger("my_app")
|
||||
session = AdvancedSQLiteSession(
|
||||
session_id="session_456",
|
||||
create_tables=True,
|
||||
logger=logger
|
||||
)
|
||||
```
|
||||
|
||||
### パラメーター
|
||||
|
||||
- `session_id` (str): 会話セッションの一意の識別子
|
||||
- `db_path` (str | Path): SQLite データベースファイルへのパス。デフォルトはインメモリストレージ用の `:memory:` です
|
||||
- `create_tables` (bool): 高度なテーブルを自動的に作成するかどうか。デフォルトは `False` です
|
||||
- `logger` (logging.Logger | None): セッション用のカスタムロガー。デフォルトはモジュールロガーです
|
||||
|
||||
## 使用状況の追跡
|
||||
|
||||
AdvancedSQLiteSession は、会話ターンごとにトークン使用状況データを保存することで、詳細な使用状況分析を提供します。 **これは、各エージェント実行後に `store_run_usage` メソッドが呼び出されることに完全に依存します。**
|
||||
|
||||
### 使用状況データの保存
|
||||
|
||||
```python
|
||||
# After each agent run, store the usage data
|
||||
result = await Runner.run(agent, "Hello", session=session)
|
||||
await session.store_run_usage(result)
|
||||
|
||||
# This stores:
|
||||
# - Total tokens used
|
||||
# - Input/output token breakdown
|
||||
# - Request count
|
||||
# - Detailed JSON token information (if available)
|
||||
```
|
||||
|
||||
### 使用状況統計の取得
|
||||
|
||||
```python
|
||||
# Get session-level usage (all branches)
|
||||
session_usage = await session.get_session_usage()
|
||||
if session_usage:
|
||||
print(f"Total requests: {session_usage['requests']}")
|
||||
print(f"Total tokens: {session_usage['total_tokens']}")
|
||||
print(f"Input tokens: {session_usage['input_tokens']}")
|
||||
print(f"Output tokens: {session_usage['output_tokens']}")
|
||||
print(f"Total turns: {session_usage['total_turns']}")
|
||||
|
||||
# Get usage for specific branch
|
||||
branch_usage = await session.get_session_usage(branch_id="main")
|
||||
|
||||
# Get usage by turn
|
||||
turn_usage = await session.get_turn_usage()
|
||||
for turn_data in turn_usage:
|
||||
print(f"Turn {turn_data['user_turn_number']}: {turn_data['total_tokens']} tokens")
|
||||
if turn_data['input_tokens_details']:
|
||||
print(f" Input details: {turn_data['input_tokens_details']}")
|
||||
if turn_data['output_tokens_details']:
|
||||
print(f" Output details: {turn_data['output_tokens_details']}")
|
||||
|
||||
# Get usage for specific turn
|
||||
turn_2_usage = await session.get_turn_usage(user_turn_number=2)
|
||||
```
|
||||
|
||||
## 会話の分岐
|
||||
|
||||
AdvancedSQLiteSession の主要機能の 1 つは、任意のユーザーメッセージから会話ブランチを作成し、別の会話パスを探索できることです。
|
||||
|
||||
### ブランチの作成
|
||||
|
||||
```python
|
||||
# Get available turns for branching
|
||||
turns = await session.get_conversation_turns()
|
||||
for turn in turns:
|
||||
print(f"Turn {turn['turn']}: {turn['content']}")
|
||||
print(f"Can branch: {turn['can_branch']}")
|
||||
|
||||
# Create a branch from turn 2
|
||||
branch_id = await session.create_branch_from_turn(2)
|
||||
print(f"Created branch: {branch_id}")
|
||||
|
||||
# Create a branch with custom name
|
||||
branch_id = await session.create_branch_from_turn(
|
||||
2,
|
||||
branch_name="alternative_path"
|
||||
)
|
||||
|
||||
# Create branch by searching for content
|
||||
branch_id = await session.create_branch_from_content(
|
||||
"weather",
|
||||
branch_name="weather_focus"
|
||||
)
|
||||
```
|
||||
|
||||
### ブランチ管理
|
||||
|
||||
```python
|
||||
# List all branches
|
||||
branches = await session.list_branches()
|
||||
for branch in branches:
|
||||
current = " (current)" if branch["is_current"] else ""
|
||||
print(f"{branch['branch_id']}: {branch['user_turns']} turns, {branch['message_count']} messages{current}")
|
||||
|
||||
# Switch between branches
|
||||
await session.switch_to_branch("main")
|
||||
await session.switch_to_branch(branch_id)
|
||||
|
||||
# Delete a branch
|
||||
await session.delete_branch(branch_id, force=True) # force=True allows deleting current branch
|
||||
```
|
||||
|
||||
### ブランチワークフロー例
|
||||
|
||||
```python
|
||||
# Original conversation
|
||||
result = await Runner.run(agent, "What's the capital of France?", session=session)
|
||||
await session.store_run_usage(result)
|
||||
|
||||
result = await Runner.run(agent, "What's the weather like there?", session=session)
|
||||
await session.store_run_usage(result)
|
||||
|
||||
# Create branch from turn 2 (weather question)
|
||||
branch_id = await session.create_branch_from_turn(2, "weather_focus")
|
||||
|
||||
# Continue in new branch with different question
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
"What are the main tourist attractions in Paris?",
|
||||
session=session
|
||||
)
|
||||
await session.store_run_usage(result)
|
||||
|
||||
# Switch back to main branch
|
||||
await session.switch_to_branch("main")
|
||||
|
||||
# Continue original conversation
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
"How expensive is it to visit?",
|
||||
session=session
|
||||
)
|
||||
await session.store_run_usage(result)
|
||||
```
|
||||
|
||||
## 構造化クエリ
|
||||
|
||||
AdvancedSQLiteSession は、会話の構造と内容を分析するための複数のメソッドを提供します。
|
||||
|
||||
### 会話分析
|
||||
|
||||
```python
|
||||
# Get conversation organized by turns
|
||||
conversation_by_turns = await session.get_conversation_by_turns()
|
||||
for turn_num, items in conversation_by_turns.items():
|
||||
print(f"Turn {turn_num}: {len(items)} items")
|
||||
for item in items:
|
||||
if item["tool_name"]:
|
||||
print(f" - {item['type']} (tool: {item['tool_name']})")
|
||||
else:
|
||||
print(f" - {item['type']}")
|
||||
|
||||
# Get tool usage statistics
|
||||
tool_usage = await session.get_tool_usage()
|
||||
for tool_name, count, turn in tool_usage:
|
||||
print(f"{tool_name}: used {count} times in turn {turn}")
|
||||
|
||||
# Find turns by content
|
||||
matching_turns = await session.find_turns_by_content("weather")
|
||||
for turn in matching_turns:
|
||||
print(f"Turn {turn['turn']}: {turn['content']}")
|
||||
```
|
||||
|
||||
### メッセージ構造
|
||||
|
||||
セッションは、次を含むメッセージ構造を自動的に追跡します。
|
||||
|
||||
- メッセージタイプ(ユーザー、assistant、tool_call など)
|
||||
- ツール呼び出しのツール名
|
||||
- ターン番号とシーケンス番号
|
||||
- ブランチとの関連付け
|
||||
- タイムスタンプ
|
||||
|
||||
## データベーススキーマ
|
||||
|
||||
AdvancedSQLiteSession は、基本的な SQLite スキーマを 2 つの追加テーブルで拡張します。
|
||||
|
||||
### message_structure テーブル
|
||||
|
||||
```sql
|
||||
CREATE TABLE message_structure (
|
||||
id INTEGER PRIMARY KEY AUTOINCREMENT,
|
||||
session_id TEXT NOT NULL,
|
||||
message_id INTEGER NOT NULL,
|
||||
branch_id TEXT NOT NULL DEFAULT 'main',
|
||||
message_type TEXT NOT NULL,
|
||||
sequence_number INTEGER NOT NULL,
|
||||
user_turn_number INTEGER,
|
||||
branch_turn_number INTEGER,
|
||||
tool_name TEXT,
|
||||
created_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP,
|
||||
FOREIGN KEY (session_id) REFERENCES agent_sessions(session_id) ON DELETE CASCADE,
|
||||
FOREIGN KEY (message_id) REFERENCES agent_messages(id) ON DELETE CASCADE
|
||||
);
|
||||
```
|
||||
|
||||
### turn_usage テーブル
|
||||
|
||||
```sql
|
||||
CREATE TABLE turn_usage (
|
||||
id INTEGER PRIMARY KEY AUTOINCREMENT,
|
||||
session_id TEXT NOT NULL,
|
||||
branch_id TEXT NOT NULL DEFAULT 'main',
|
||||
user_turn_number INTEGER NOT NULL,
|
||||
requests INTEGER DEFAULT 0,
|
||||
input_tokens INTEGER DEFAULT 0,
|
||||
output_tokens INTEGER DEFAULT 0,
|
||||
total_tokens INTEGER DEFAULT 0,
|
||||
input_tokens_details JSON,
|
||||
output_tokens_details JSON,
|
||||
created_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP,
|
||||
FOREIGN KEY (session_id) REFERENCES agent_sessions(session_id) ON DELETE CASCADE,
|
||||
UNIQUE(session_id, branch_id, user_turn_number)
|
||||
);
|
||||
```
|
||||
|
||||
## 完全な例
|
||||
|
||||
すべての機能を包括的に示す [完全な例](https://github.com/openai/openai-agents-python/tree/main/examples/memory/advanced_sqlite_session_example.py) を確認してください。
|
||||
|
||||
|
||||
## API リファレンス
|
||||
|
||||
- [`AdvancedSQLiteSession`][agents.extensions.memory.advanced_sqlite_session.AdvancedSQLiteSession] - メインクラス
|
||||
- [`Session`][agents.memory.session.Session] - ベースセッションプロトコル
|
||||
@@ -0,0 +1,179 @@
|
||||
---
|
||||
search:
|
||||
exclude: true
|
||||
---
|
||||
# 暗号化セッション
|
||||
|
||||
`EncryptedSession` は、任意のセッション実装に透過的な暗号化を提供し、古いアイテムの自動期限切れによって会話データを保護します。
|
||||
|
||||
## 機能
|
||||
|
||||
- **透過的な暗号化**: 任意のセッションを Fernet 暗号化でラップします
|
||||
- **セッションごとのキー**: HKDF キー導出を使用して、セッションごとに一意の暗号化を行います
|
||||
- **自動期限切れ**: TTL が期限切れになると、古いアイテムは取得時に黙ってスキップされます
|
||||
- **ドロップイン置換**: 既存の任意のセッション実装で動作します
|
||||
|
||||
## インストール
|
||||
|
||||
暗号化セッションには `encrypt` extra が必要です。
|
||||
|
||||
```bash
|
||||
pip install openai-agents[encrypt]
|
||||
```
|
||||
|
||||
## クイックスタート
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
from agents import Agent, Runner
|
||||
from agents.extensions.memory import EncryptedSession, SQLAlchemySession
|
||||
|
||||
async def main():
|
||||
agent = Agent("Assistant")
|
||||
|
||||
# Create underlying session
|
||||
underlying_session = SQLAlchemySession.from_url(
|
||||
"user-123",
|
||||
url="sqlite+aiosqlite:///:memory:",
|
||||
create_tables=True
|
||||
)
|
||||
|
||||
# Wrap with encryption
|
||||
session = EncryptedSession(
|
||||
session_id="user-123",
|
||||
underlying_session=underlying_session,
|
||||
encryption_key="your-secret-key-here",
|
||||
ttl=600 # 10 minutes
|
||||
)
|
||||
|
||||
result = await Runner.run(agent, "Hello", session=session)
|
||||
print(result.final_output)
|
||||
|
||||
if __name__ == "__main__":
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
## 設定
|
||||
|
||||
### 暗号化キー
|
||||
|
||||
暗号化キーには、Fernet キーまたは任意の文字列を指定できます。
|
||||
|
||||
```python
|
||||
from agents.extensions.memory import EncryptedSession
|
||||
|
||||
# Using a Fernet key (base64-encoded)
|
||||
session = EncryptedSession(
|
||||
session_id="user-123",
|
||||
underlying_session=underlying_session,
|
||||
encryption_key="your-fernet-key-here",
|
||||
ttl=600
|
||||
)
|
||||
|
||||
# Using a raw string (will be derived to a key)
|
||||
session = EncryptedSession(
|
||||
session_id="user-123",
|
||||
underlying_session=underlying_session,
|
||||
encryption_key="my-secret-password",
|
||||
ttl=600
|
||||
)
|
||||
```
|
||||
|
||||
### TTL (有効期間)
|
||||
|
||||
暗号化されたアイテムが有効であり続ける期間を設定します。
|
||||
|
||||
```python
|
||||
# Items expire after 1 hour
|
||||
session = EncryptedSession(
|
||||
session_id="user-123",
|
||||
underlying_session=underlying_session,
|
||||
encryption_key="secret",
|
||||
ttl=3600 # 1 hour in seconds
|
||||
)
|
||||
|
||||
# Items expire after 1 day
|
||||
session = EncryptedSession(
|
||||
session_id="user-123",
|
||||
underlying_session=underlying_session,
|
||||
encryption_key="secret",
|
||||
ttl=86400 # 24 hours in seconds
|
||||
)
|
||||
```
|
||||
|
||||
## さまざまなセッションタイプでの使用
|
||||
|
||||
### SQLite セッションでの使用
|
||||
|
||||
```python
|
||||
from agents import SQLiteSession
|
||||
from agents.extensions.memory import EncryptedSession
|
||||
|
||||
# Create encrypted SQLite session
|
||||
underlying = SQLiteSession("user-123", "conversations.db")
|
||||
|
||||
session = EncryptedSession(
|
||||
session_id="user-123",
|
||||
underlying_session=underlying,
|
||||
encryption_key="secret-key"
|
||||
)
|
||||
```
|
||||
|
||||
### SQLAlchemy セッションでの使用
|
||||
|
||||
```python
|
||||
from agents.extensions.memory import EncryptedSession, SQLAlchemySession
|
||||
|
||||
# Create encrypted SQLAlchemy session
|
||||
underlying = SQLAlchemySession.from_url(
|
||||
"user-123",
|
||||
url="postgresql+asyncpg://user:pass@localhost/db",
|
||||
create_tables=True
|
||||
)
|
||||
|
||||
session = EncryptedSession(
|
||||
session_id="user-123",
|
||||
underlying_session=underlying,
|
||||
encryption_key="secret-key"
|
||||
)
|
||||
```
|
||||
|
||||
!!! warning "高度なセッション機能"
|
||||
|
||||
`EncryptedSession` を `AdvancedSQLiteSession` のような高度なセッション実装と使用する場合は、次の点に注意してください。
|
||||
|
||||
- メッセージ内容が暗号化されるため、`find_turns_by_content()` のようなメソッドは効果的に動作しません
|
||||
- コンテンツベースの検索は暗号化されたデータに対して実行されるため、有効性が制限されます
|
||||
|
||||
|
||||
|
||||
## キー導出
|
||||
|
||||
EncryptedSession は HKDF (HMAC-based Key Derivation Function) を使用して、セッションごとに一意の暗号化キーを導出します。
|
||||
|
||||
- **マスターキー**: 提供された暗号化キー
|
||||
- **セッションソルト**: セッション ID
|
||||
- **情報文字列**: `"agents.session-store.hkdf.v1"`
|
||||
- **出力**: 32 バイトの Fernet キー
|
||||
|
||||
これにより、次のことが保証されます。
|
||||
- 各セッションに一意の暗号化キーがあります
|
||||
- マスターキーがなければキーを導出できません
|
||||
- セッションデータを異なるセッション間で復号できません
|
||||
|
||||
## 自動期限切れ
|
||||
|
||||
アイテムが TTL を超えると、取得時に自動的にスキップされます。
|
||||
|
||||
```python
|
||||
# Items older than TTL are silently ignored
|
||||
items = await session.get_items() # Only returns non-expired items
|
||||
|
||||
# Expired items don't affect session behavior
|
||||
result = await Runner.run(agent, "Continue conversation", session=session)
|
||||
```
|
||||
|
||||
## API リファレンス
|
||||
|
||||
- [`EncryptedSession`][agents.extensions.memory.encrypt_session.EncryptedSession] - メインクラス
|
||||
- [`Session`][agents.memory.session.Session] - ベースセッションプロトコル
|
||||
@@ -0,0 +1,711 @@
|
||||
---
|
||||
search:
|
||||
exclude: true
|
||||
---
|
||||
# セッション
|
||||
|
||||
Agents SDK は、複数のエージェント実行にまたがって会話履歴を自動的に維持する組み込みのセッションメモリを提供し、ターン間で `.to_input_list()` を手動で扱う必要をなくします。
|
||||
|
||||
セッションは特定のセッションの会話履歴を保存し、明示的な手動メモリ管理を必要とせずに、エージェントがコンテキストを維持できるようにします。これは、エージェントに以前のやり取りを覚えておいてほしいチャットアプリケーションや複数ターンの会話を構築する場合に特に便利です。
|
||||
|
||||
SDK にクライアント側メモリを管理させたい場合は、セッションを使用します。セッションは、同じ実行内で `conversation_id`、`previous_response_id`、または `auto_previous_response_id` と組み合わせることはできません。代わりに OpenAI のサーバー管理による継続を使用したい場合は、セッションを重ねて使うのではなく、それらの仕組みのいずれかを選択してください。
|
||||
|
||||
## クイックスタート
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner, SQLiteSession
|
||||
|
||||
# Create agent
|
||||
agent = Agent(
|
||||
name="Assistant",
|
||||
instructions="Reply very concisely.",
|
||||
)
|
||||
|
||||
# Create a session instance with a session ID
|
||||
session = SQLiteSession("conversation_123")
|
||||
|
||||
# First turn
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
"What city is the Golden Gate Bridge in?",
|
||||
session=session
|
||||
)
|
||||
print(result.final_output) # "San Francisco"
|
||||
|
||||
# Second turn - agent automatically remembers previous context
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
"What state is it in?",
|
||||
session=session
|
||||
)
|
||||
print(result.final_output) # "California"
|
||||
|
||||
# Also works with synchronous runner
|
||||
result = Runner.run_sync(
|
||||
agent,
|
||||
"What's the population?",
|
||||
session=session
|
||||
)
|
||||
print(result.final_output) # "Approximately 39 million"
|
||||
```
|
||||
|
||||
## 同じセッションによる中断された実行の再開
|
||||
|
||||
実行が承認待ちで一時停止した場合は、同じセッションインスタンス(または同じバッキングストアを指す別のセッションインスタンス)で再開し、再開されたターンが同じ保存済み会話履歴を継続するようにします。
|
||||
|
||||
```python
|
||||
result = await Runner.run(agent, "Delete temporary files that are no longer needed.", session=session)
|
||||
|
||||
if result.interruptions:
|
||||
state = result.to_state()
|
||||
for interruption in result.interruptions:
|
||||
state.approve(interruption)
|
||||
result = await Runner.run(agent, state, session=session)
|
||||
```
|
||||
|
||||
## コアセッション動作
|
||||
|
||||
セッションメモリが有効な場合:
|
||||
|
||||
1. **各実行の前**: ランナーはセッションの会話履歴を自動的に取得し、入力アイテムの前に追加します。
|
||||
2. **各実行の後**: 実行中に生成されたすべての新しいアイテム(ユーザー入力、アシスタントの応答、ツール呼び出しなど)がセッションに自動的に保存されます。
|
||||
3. **コンテキストの保持**: 同じセッションでの後続の各実行には完全な会話履歴が含まれるため、エージェントはコンテキストを維持できます。
|
||||
|
||||
これにより、`.to_input_list()` を手動で呼び出したり、実行間の会話状態を管理したりする必要がなくなります。
|
||||
|
||||
## 履歴と新しい入力のマージ方法の制御
|
||||
|
||||
セッションを渡すと、ランナーは通常、モデル入力を次のように準備します。
|
||||
|
||||
1. セッション履歴(`session.get_items(...)` から取得)
|
||||
2. 新しいターン入力
|
||||
|
||||
モデル呼び出しの前にこのマージ手順をカスタマイズするには、[`RunConfig.session_input_callback`][agents.run.RunConfig.session_input_callback] を使用します。コールバックは次の 2 つのリストを受け取ります。
|
||||
|
||||
- `history`: 取得されたセッション履歴(すでに入力アイテム形式に正規化済み)
|
||||
- `new_input`: 現在のターンの新しい入力アイテム
|
||||
|
||||
モデルに送信する最終的な入力アイテムのリストを返します。
|
||||
|
||||
コールバックは両方のリストのコピーを受け取るため、安全に変更できます。返されたリストはそのターンのモデル入力を制御しますが、SDK は新しいターンに属するアイテムのみを永続化します。そのため、古い履歴を並べ替えたりフィルタリングしたりしても、古いセッションアイテムが新しい入力として再度保存されることはありません。
|
||||
|
||||
```python
|
||||
from agents import Agent, RunConfig, Runner, SQLiteSession
|
||||
|
||||
|
||||
def keep_recent_history(history, new_input):
|
||||
# Keep only the last 10 history items, then append the new turn.
|
||||
return history[-10:] + new_input
|
||||
|
||||
|
||||
agent = Agent(name="Assistant")
|
||||
session = SQLiteSession("conversation_123")
|
||||
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
"Continue from the latest updates only.",
|
||||
session=session,
|
||||
run_config=RunConfig(session_input_callback=keep_recent_history),
|
||||
)
|
||||
```
|
||||
|
||||
セッションがアイテムを保存する方法を変更せずに、カスタムの枝刈り、並べ替え、または履歴の選択的な取り込みが必要な場合に使用します。モデル呼び出しの直前にさらに最終的な処理が必要な場合は、[エージェント実行ガイド](../running_agents.md)の [`call_model_input_filter`][agents.run.RunConfig.call_model_input_filter] を使用してください。
|
||||
|
||||
## 取得する履歴の制限
|
||||
|
||||
各実行の前に取得する履歴の量を制御するには、[`SessionSettings`][agents.memory.SessionSettings] を使用します。
|
||||
|
||||
- `SessionSettings(limit=None)`(デフォルト): 利用可能なすべてのセッションアイテムを取得します
|
||||
- `SessionSettings(limit=N)`: 直近の `N` アイテムのみを取得します
|
||||
|
||||
これは、[`RunConfig.session_settings`][agents.run.RunConfig.session_settings] を介して実行ごとに適用できます。
|
||||
|
||||
```python
|
||||
from agents import Agent, RunConfig, Runner, SessionSettings, SQLiteSession
|
||||
|
||||
agent = Agent(name="Assistant")
|
||||
session = SQLiteSession("conversation_123")
|
||||
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
"Summarize our recent discussion.",
|
||||
session=session,
|
||||
run_config=RunConfig(session_settings=SessionSettings(limit=50)),
|
||||
)
|
||||
```
|
||||
|
||||
セッション実装がデフォルトのセッション設定を公開している場合、`RunConfig.session_settings` はその実行について `None` ではない値を上書きします。これは、セッションのデフォルト動作を変更せずに取得サイズを上限設定したい長い会話で便利です。
|
||||
|
||||
## メモリ操作
|
||||
|
||||
### 基本操作
|
||||
|
||||
セッションは、会話履歴を管理するためのいくつかの操作をサポートしています。
|
||||
|
||||
```python
|
||||
from agents import SQLiteSession
|
||||
|
||||
session = SQLiteSession("user_123", "conversations.db")
|
||||
|
||||
# Get all items in a session
|
||||
items = await session.get_items()
|
||||
|
||||
# Add new items to a session
|
||||
new_items = [
|
||||
{"role": "user", "content": "Hello"},
|
||||
{"role": "assistant", "content": "Hi there!"}
|
||||
]
|
||||
await session.add_items(new_items)
|
||||
|
||||
# Remove and return the most recent item
|
||||
last_item = await session.pop_item()
|
||||
print(last_item) # {"role": "assistant", "content": "Hi there!"}
|
||||
|
||||
# Clear all items from a session
|
||||
await session.clear_session()
|
||||
```
|
||||
|
||||
### 修正のための pop_item の使用
|
||||
|
||||
`pop_item` メソッドは、会話内の最後のアイテムを取り消したり変更したりしたい場合に特に便利です。
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner, SQLiteSession
|
||||
|
||||
agent = Agent(name="Assistant")
|
||||
session = SQLiteSession("correction_example")
|
||||
|
||||
# Initial conversation
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
"What's 2 + 2?",
|
||||
session=session
|
||||
)
|
||||
print(f"Agent: {result.final_output}")
|
||||
|
||||
# User wants to correct their question
|
||||
assistant_item = await session.pop_item() # Remove agent's response
|
||||
user_item = await session.pop_item() # Remove user's question
|
||||
|
||||
# Ask a corrected question
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
"What's 2 + 3?",
|
||||
session=session
|
||||
)
|
||||
print(f"Agent: {result.final_output}")
|
||||
```
|
||||
|
||||
## 組み込みセッション実装
|
||||
|
||||
SDK は、さまざまなユースケース向けに複数のセッション実装を提供しています。
|
||||
|
||||
### 組み込みセッション実装の選択
|
||||
|
||||
以下の詳細な例を読む前に、開始点を選ぶためにこの表を使用してください。
|
||||
|
||||
| セッションタイプ | 最適な用途 | 注記 |
|
||||
| --- | --- | --- |
|
||||
| `SQLiteSession` | ローカル開発とシンプルなアプリ | 組み込み、軽量、ファイルバックまたはインメモリ |
|
||||
| `AsyncSQLiteSession` | `aiosqlite` を使用した非同期 SQLite | 非同期ドライバー対応の拡張バックエンド |
|
||||
| `RedisSession` | ワーカーやサービス間で共有するメモリ | 低レイテンシの分散デプロイに適しています |
|
||||
| `SQLAlchemySession` | 既存データベースを使用する本番アプリ | SQLAlchemy がサポートするデータベースで動作します |
|
||||
| `MongoDBSession` | すでに MongoDB を使用しているアプリ、またはマルチプロセスストレージが必要なアプリ | 非同期 pymongo;順序付け用のアトミックシーケンスカウンター |
|
||||
| `DaprSession` | Dapr サイドカーを使用するクラウドネイティブデプロイ | 複数のステートストアに加え、TTL と整合性制御をサポートします |
|
||||
| `OpenAIConversationsSession` | OpenAI でのサーバー管理ストレージ | OpenAI Conversations API をバックエンドとする履歴 |
|
||||
| `OpenAIResponsesCompactionSession` | 自動圧縮を伴う長い会話 | 別のセッションバックエンドをラップします |
|
||||
| `AdvancedSQLiteSession` | SQLite に加えて分岐や分析 | より多機能です。専用ページを参照してください |
|
||||
| `EncryptedSession` | 別のセッション上での暗号化と TTL | ラッパーです。まず基盤となるバックエンドを選択してください |
|
||||
|
||||
一部の実装には、追加の詳細を含む専用ページがあります。それらは各サブセクション内でリンクされています。
|
||||
|
||||
ChatKit 用の Python サーバーを実装している場合は、ChatKit のスレッドとアイテムの永続化に `chatkit.store.Store` 実装を使用してください。`SQLAlchemySession` などの Agents SDK セッションは SDK 側の会話履歴を管理しますが、ChatKit のストアのドロップイン置き換えではありません。[ChatKit データストアの実装に関する `chatkit-python` ガイド](https://github.com/openai/chatkit-python/blob/main/docs/guides/respond-to-user-message.md#implement-your-chatkit-data-store)を参照してください。
|
||||
|
||||
### OpenAI Conversations API セッション
|
||||
|
||||
`OpenAIConversationsSession` を通じて [OpenAI の Conversations API](https://platform.openai.com/docs/api-reference/conversations)を使用します。
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner, OpenAIConversationsSession
|
||||
|
||||
# Create agent
|
||||
agent = Agent(
|
||||
name="Assistant",
|
||||
instructions="Reply very concisely.",
|
||||
)
|
||||
|
||||
# Create a new conversation
|
||||
session = OpenAIConversationsSession()
|
||||
|
||||
# Optionally resume a previous conversation by passing a conversation ID
|
||||
# session = OpenAIConversationsSession(conversation_id="conv_123")
|
||||
|
||||
# Start conversation
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
"What city is the Golden Gate Bridge in?",
|
||||
session=session
|
||||
)
|
||||
print(result.final_output) # "San Francisco"
|
||||
|
||||
# Continue the conversation
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
"What state is it in?",
|
||||
session=session
|
||||
)
|
||||
print(result.final_output) # "California"
|
||||
```
|
||||
|
||||
### OpenAI Responses 圧縮セッション
|
||||
|
||||
Responses API(`responses.compact`)で保存済みの会話履歴を圧縮するには、`OpenAIResponsesCompactionSession` を使用します。これは基盤となるセッションをラップし、`should_trigger_compaction` に基づいて各ターンの後に自動的に圧縮できます。`OpenAIConversationsSession` をこれでラップしないでください。この 2 つの機能は異なる方法で履歴を管理します。
|
||||
|
||||
#### 一般的な使用方法(自動圧縮)
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner, SQLiteSession
|
||||
from agents.memory import OpenAIResponsesCompactionSession
|
||||
|
||||
underlying = SQLiteSession("conversation_123")
|
||||
session = OpenAIResponsesCompactionSession(
|
||||
session_id="conversation_123",
|
||||
underlying_session=underlying,
|
||||
)
|
||||
|
||||
agent = Agent(name="Assistant")
|
||||
result = await Runner.run(agent, "Hello", session=session)
|
||||
print(result.final_output)
|
||||
```
|
||||
|
||||
デフォルトでは、候補しきい値に達すると各ターンの後に圧縮が実行されます。
|
||||
|
||||
`compaction_mode="previous_response_id"` は、Responses API の応答 ID でターンをすでに連鎖させている場合に最も適しています。`compaction_mode="input"` は、代わりに現在のセッションアイテムから圧縮リクエストを再構築します。これは、応答チェーンが利用できない場合や、セッション内容を信頼できる情報源にしたい場合に便利です。デフォルトの `"auto"` は、利用可能な中で最も安全な選択肢を選びます。
|
||||
|
||||
エージェントが `ModelSettings(store=False)` で実行される場合、Responses API は後で検索するための最後の応答を保持しません。このステートレスな構成では、デフォルトの `"auto"` モードは `previous_response_id` に依存するのではなく、入力ベースの圧縮にフォールバックします。完全な例については、[`examples/memory/compaction_session_stateless_example.py`](https://github.com/openai/openai-agents-python/tree/main/examples/memory/compaction_session_stateless_example.py) を参照してください。
|
||||
|
||||
#### auto-compaction によるストリーミングのブロック
|
||||
|
||||
圧縮はセッション履歴をクリアして書き換えるため、SDK は実行完了とみなす前に圧縮の完了を待ちます。ストリーミングモードでは、圧縮が重い場合、最後の出力トークンの後も `run.stream_events()` が数秒間開いたままになることがあります。
|
||||
|
||||
低レイテンシのストリーミングや高速なターン処理が必要な場合は、自動圧縮を無効にし、ターン間(またはアイドル時間中)に自分で `run_compaction()` を呼び出してください。独自の基準に基づいて、いつ圧縮を強制するかを決めることができます。
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner, SQLiteSession
|
||||
from agents.memory import OpenAIResponsesCompactionSession
|
||||
|
||||
underlying = SQLiteSession("conversation_123")
|
||||
session = OpenAIResponsesCompactionSession(
|
||||
session_id="conversation_123",
|
||||
underlying_session=underlying,
|
||||
# Disable triggering the auto compaction
|
||||
should_trigger_compaction=lambda _: False,
|
||||
)
|
||||
|
||||
agent = Agent(name="Assistant")
|
||||
result = await Runner.run(agent, "Hello", session=session)
|
||||
|
||||
# Decide when to compact (e.g., on idle, every N turns, or size thresholds).
|
||||
await session.run_compaction({"force": True})
|
||||
```
|
||||
|
||||
### SQLite セッション
|
||||
|
||||
SQLite を使用するデフォルトの軽量セッション実装です。
|
||||
|
||||
```python
|
||||
from agents import SQLiteSession
|
||||
|
||||
# In-memory database (lost when process ends)
|
||||
session = SQLiteSession("user_123")
|
||||
|
||||
# Persistent file-based database
|
||||
session = SQLiteSession("user_123", "conversations.db")
|
||||
|
||||
# Use the session
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
"Hello",
|
||||
session=session
|
||||
)
|
||||
```
|
||||
|
||||
### 非同期 SQLite セッション
|
||||
|
||||
`aiosqlite` をバックエンドとする SQLite の永続化が必要な場合は、`AsyncSQLiteSession` を使用します。
|
||||
|
||||
```bash
|
||||
pip install aiosqlite
|
||||
```
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner
|
||||
from agents.extensions.memory import AsyncSQLiteSession
|
||||
|
||||
agent = Agent(name="Assistant")
|
||||
session = AsyncSQLiteSession("user_123", db_path="conversations.db")
|
||||
result = await Runner.run(agent, "Hello", session=session)
|
||||
```
|
||||
|
||||
### Redis セッション
|
||||
|
||||
複数のワーカーまたはサービス間で共有セッションメモリを使用するには、`RedisSession` を使用します。
|
||||
|
||||
```bash
|
||||
pip install openai-agents[redis]
|
||||
```
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner
|
||||
from agents.extensions.memory import RedisSession
|
||||
|
||||
agent = Agent(name="Assistant")
|
||||
session = RedisSession.from_url(
|
||||
"user_123",
|
||||
url="redis://localhost:6379/0",
|
||||
)
|
||||
result = await Runner.run(agent, "Hello", session=session)
|
||||
```
|
||||
|
||||
### SQLAlchemy セッション
|
||||
|
||||
SQLAlchemy がサポートする任意のデータベースを使用した、本番対応の Agents SDK セッション永続化です。
|
||||
|
||||
```python
|
||||
from agents.extensions.memory import SQLAlchemySession
|
||||
|
||||
# Using database URL
|
||||
session = SQLAlchemySession.from_url(
|
||||
"user_123",
|
||||
url="postgresql+asyncpg://user:pass@localhost/db",
|
||||
create_tables=True
|
||||
)
|
||||
|
||||
# Using existing engine
|
||||
from sqlalchemy.ext.asyncio import create_async_engine
|
||||
engine = create_async_engine("postgresql+asyncpg://user:pass@localhost/db")
|
||||
session = SQLAlchemySession("user_123", engine=engine, create_tables=True)
|
||||
```
|
||||
|
||||
詳細なドキュメントについては、[SQLAlchemy セッション](sqlalchemy_session.md)を参照してください。
|
||||
|
||||
### Dapr セッション
|
||||
|
||||
すでに Dapr サイドカーを実行している場合、またはエージェントコードを変更せずに異なるステートストアバックエンドへ移行できるセッションストレージが必要な場合は、`DaprSession` を使用します。
|
||||
|
||||
```bash
|
||||
pip install openai-agents[dapr]
|
||||
```
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner
|
||||
from agents.extensions.memory import DaprSession
|
||||
|
||||
agent = Agent(name="Assistant")
|
||||
|
||||
async with DaprSession.from_address(
|
||||
"user_123",
|
||||
state_store_name="statestore",
|
||||
dapr_address="localhost:50001",
|
||||
) as session:
|
||||
result = await Runner.run(agent, "Hello", session=session)
|
||||
print(result.final_output)
|
||||
```
|
||||
|
||||
注記:
|
||||
|
||||
- `from_address(...)` は Dapr クライアントを作成し、所有します。アプリがすでにクライアントを管理している場合は、`dapr_client=...` を指定して `DaprSession(...)` を直接構築してください。
|
||||
- 基盤となるステートストアが TTL をサポートしている場合に古いセッションデータを自動的に期限切れにするには、`ttl=...` を渡します。
|
||||
- より強い read-after-write 保証が必要な場合は、`consistency=DAPR_CONSISTENCY_STRONG` を渡します。
|
||||
- Dapr Python SDK は HTTP サイドカーエンドポイントもチェックします。ローカル開発では、`dapr_address` で使用する gRPC ポートに加えて、`--dapr-http-port 3500` でも Dapr を起動してください。
|
||||
- ローカルコンポーネントやトラブルシューティングを含む完全なセットアップ手順については、[`examples/memory/dapr_session_example.py`](https://github.com/openai/openai-agents-python/tree/main/examples/memory/dapr_session_example.py) を参照してください。
|
||||
|
||||
|
||||
### MongoDB セッション
|
||||
|
||||
すでに MongoDB を使用しているアプリケーション、または水平スケーラブルでマルチプロセス対応のセッションストレージが必要なアプリケーションには、`MongoDBSession` を使用します。
|
||||
|
||||
```bash
|
||||
pip install openai-agents[mongodb]
|
||||
```
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner
|
||||
from agents.extensions.memory import MongoDBSession
|
||||
|
||||
agent = Agent(name="Assistant")
|
||||
|
||||
# Create from URI — owns the client and closes it when session.close() is called
|
||||
session = MongoDBSession.from_uri(
|
||||
"user-123",
|
||||
uri="mongodb://localhost:27017",
|
||||
database="agents",
|
||||
)
|
||||
result = await Runner.run(agent, "Hello", session=session)
|
||||
print(result.final_output)
|
||||
await session.close()
|
||||
```
|
||||
|
||||
注記:
|
||||
|
||||
- `from_uri(...)` は `AsyncMongoClient` を作成し、所有し、`session.close()` で閉じます。アプリケーションがすでにクライアントを管理している場合は、`client=...` を指定して `MongoDBSession(...)` を直接構築してください。その場合、`session.close()` は no-op となり、ライフサイクルは呼び出し元が保持します。
|
||||
- ほかの変更なしに、`from_uri(...)` に `mongodb+srv://user:password@cluster.example.mongodb.net` URI を渡すことで [MongoDB Atlas](https://www.mongodb.com/products/platform) に接続できます。
|
||||
- 2 つのコレクションが使用され、どちらの名前も `sessions_collection=`(デフォルトは `agent_sessions`)と `messages_collection=`(デフォルトは `agent_messages`)で設定できます。インデックスは初回使用時に自動的に作成されます。各メッセージドキュメントは、同時実行の書き込み元やプロセスをまたいで順序を保持する単調増加の `seq` カウンターを持ちます。
|
||||
- 最初の実行前に接続性を確認するには、`await session.ping()` を使用します。
|
||||
|
||||
### 高度な SQLite セッション
|
||||
|
||||
会話の分岐、使用状況分析、構造化クエリを備えた拡張 SQLite セッションです。
|
||||
|
||||
```python
|
||||
from agents.extensions.memory import AdvancedSQLiteSession
|
||||
|
||||
# Create with advanced features
|
||||
session = AdvancedSQLiteSession(
|
||||
session_id="user_123",
|
||||
db_path="conversations.db",
|
||||
create_tables=True
|
||||
)
|
||||
|
||||
# Automatic usage tracking
|
||||
result = await Runner.run(agent, "Hello", session=session)
|
||||
await session.store_run_usage(result) # Track token usage
|
||||
|
||||
# Conversation branching
|
||||
await session.create_branch_from_turn(2) # Branch from turn 2
|
||||
```
|
||||
|
||||
詳細なドキュメントについては、[高度な SQLite セッション](advanced_sqlite_session.md)を参照してください。
|
||||
|
||||
### 暗号化セッション
|
||||
|
||||
任意のセッション実装向けの透過的な暗号化ラッパーです。
|
||||
|
||||
```python
|
||||
from agents.extensions.memory import EncryptedSession, SQLAlchemySession
|
||||
|
||||
# Create underlying session
|
||||
underlying_session = SQLAlchemySession.from_url(
|
||||
"user_123",
|
||||
url="sqlite+aiosqlite:///conversations.db",
|
||||
create_tables=True
|
||||
)
|
||||
|
||||
# Wrap with encryption and TTL
|
||||
session = EncryptedSession(
|
||||
session_id="user_123",
|
||||
underlying_session=underlying_session,
|
||||
encryption_key="your-secret-key",
|
||||
ttl=600 # 10 minutes
|
||||
)
|
||||
|
||||
result = await Runner.run(agent, "Hello", session=session)
|
||||
```
|
||||
|
||||
詳細なドキュメントについては、[暗号化セッション](encrypted_session.md)を参照してください。
|
||||
|
||||
### その他のセッションタイプ
|
||||
|
||||
組み込みの選択肢はほかにもいくつかあります。`examples/memory/` と `extensions/memory/` 以下のソースコードを参照してください。
|
||||
|
||||
## 運用パターン
|
||||
|
||||
### セッション ID の命名
|
||||
|
||||
会話を整理しやすい、意味のあるセッション ID を使用してください。
|
||||
|
||||
- ユーザーベース: `"user_12345"`
|
||||
- スレッドベース: `"thread_abc123"`
|
||||
- コンテキストベース: `"support_ticket_456"`
|
||||
|
||||
### メモリの永続化
|
||||
|
||||
- 一時的な会話にはインメモリ SQLite(`SQLiteSession("session_id")`)を使用します
|
||||
- 永続的な会話にはファイルベース SQLite(`SQLiteSession("session_id", "path/to/db.sqlite")`)を使用します
|
||||
- `aiosqlite` ベースの実装が必要な場合は、非同期 SQLite(`AsyncSQLiteSession("session_id", db_path="...")`)を使用します
|
||||
- 共有された低レイテンシのセッションメモリには、Redis バックのセッション(`RedisSession.from_url("session_id", url="redis://...")`)を使用します
|
||||
- SQLAlchemy がサポートする既存データベースを持つ本番システムには、SQLAlchemy を利用したセッション(`SQLAlchemySession("session_id", engine=engine, create_tables=True)`) を使用します
|
||||
- すでに MongoDB を使用しているアプリケーション、またはマルチプロセスで水平スケーラブルなセッションストレージが必要なアプリケーションには、MongoDB セッション(`MongoDBSession.from_uri("session_id", uri="mongodb://localhost:27017")`)を使用します
|
||||
- 組み込みのテレメトリ、トレーシング、データ分離を備えた 30 以上のデータベースバックエンドをサポートする本番クラウドネイティブデプロイには、Dapr ステートストアセッション(`DaprSession.from_address("session_id", state_store_name="statestore", dapr_address="localhost:50001")`)を使用します
|
||||
- OpenAI Conversations API に履歴を保存したい場合は、OpenAI がホストするストレージ(`OpenAIConversationsSession()`)を使用します
|
||||
- 透過的な暗号化と TTL ベースの有効期限で任意のセッションをラップするには、暗号化セッション(`EncryptedSession(session_id, underlying_session, encryption_key)`)を使用します
|
||||
- より高度なユースケースでは、ほかの本番システム(たとえば Django)向けのカスタムセッションバックエンドの実装を検討してください
|
||||
|
||||
### 複数セッション
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner, SQLiteSession
|
||||
|
||||
agent = Agent(name="Assistant")
|
||||
|
||||
# Different sessions maintain separate conversation histories
|
||||
session_1 = SQLiteSession("user_123", "conversations.db")
|
||||
session_2 = SQLiteSession("user_456", "conversations.db")
|
||||
|
||||
result1 = await Runner.run(
|
||||
agent,
|
||||
"Help me with my account",
|
||||
session=session_1
|
||||
)
|
||||
result2 = await Runner.run(
|
||||
agent,
|
||||
"What are my charges?",
|
||||
session=session_2
|
||||
)
|
||||
```
|
||||
|
||||
### セッション共有
|
||||
|
||||
```python
|
||||
# Different agents can share the same session
|
||||
support_agent = Agent(name="Support")
|
||||
billing_agent = Agent(name="Billing")
|
||||
session = SQLiteSession("user_123")
|
||||
|
||||
# Both agents will see the same conversation history
|
||||
result1 = await Runner.run(
|
||||
support_agent,
|
||||
"Help me with my account",
|
||||
session=session
|
||||
)
|
||||
result2 = await Runner.run(
|
||||
billing_agent,
|
||||
"What are my charges?",
|
||||
session=session
|
||||
)
|
||||
```
|
||||
|
||||
## 完全な例
|
||||
|
||||
セッションメモリの動作を示す完全な例を以下に示します。
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
from agents import Agent, Runner, SQLiteSession
|
||||
|
||||
|
||||
async def main():
|
||||
# Create an agent
|
||||
agent = Agent(
|
||||
name="Assistant",
|
||||
instructions="Reply very concisely.",
|
||||
)
|
||||
|
||||
# Create a session instance that will persist across runs
|
||||
session = SQLiteSession("conversation_123", "conversation_history.db")
|
||||
|
||||
print("=== Sessions Example ===")
|
||||
print("The agent will remember previous messages automatically.\n")
|
||||
|
||||
# First turn
|
||||
print("First turn:")
|
||||
print("User: What city is the Golden Gate Bridge in?")
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
"What city is the Golden Gate Bridge in?",
|
||||
session=session
|
||||
)
|
||||
print(f"Assistant: {result.final_output}")
|
||||
print()
|
||||
|
||||
# Second turn - the agent will remember the previous conversation
|
||||
print("Second turn:")
|
||||
print("User: What state is it in?")
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
"What state is it in?",
|
||||
session=session
|
||||
)
|
||||
print(f"Assistant: {result.final_output}")
|
||||
print()
|
||||
|
||||
# Third turn - continuing the conversation
|
||||
print("Third turn:")
|
||||
print("User: What's the population of that state?")
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
"What's the population of that state?",
|
||||
session=session
|
||||
)
|
||||
print(f"Assistant: {result.final_output}")
|
||||
print()
|
||||
|
||||
print("=== Conversation Complete ===")
|
||||
print("Notice how the agent remembered the context from previous turns!")
|
||||
print("Sessions automatically handles conversation history.")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
## カスタムセッション実装
|
||||
|
||||
[`Session`][agents.memory.session.Session] プロトコルに従うクラスを作成することで、独自のセッションメモリを実装できます。
|
||||
|
||||
```python
|
||||
from agents.memory.session import SessionABC
|
||||
from agents.items import TResponseInputItem
|
||||
from typing import List
|
||||
|
||||
class MyCustomSession(SessionABC):
|
||||
"""Custom session implementation following the Session protocol."""
|
||||
|
||||
def __init__(self, session_id: str):
|
||||
self.session_id = session_id
|
||||
# Your initialization here
|
||||
|
||||
async def get_items(self, limit: int | None = None) -> List[TResponseInputItem]:
|
||||
"""Retrieve conversation history for this session."""
|
||||
# Your implementation here
|
||||
pass
|
||||
|
||||
async def add_items(self, items: List[TResponseInputItem]) -> None:
|
||||
"""Store new items for this session."""
|
||||
# Your implementation here
|
||||
pass
|
||||
|
||||
async def pop_item(self) -> TResponseInputItem | None:
|
||||
"""Remove and return the most recent item from this session."""
|
||||
# Your implementation here
|
||||
pass
|
||||
|
||||
async def clear_session(self) -> None:
|
||||
"""Clear all items for this session."""
|
||||
# Your implementation here
|
||||
pass
|
||||
|
||||
# Use your custom session
|
||||
agent = Agent(name="Assistant")
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
"Hello",
|
||||
session=MyCustomSession("my_session")
|
||||
)
|
||||
```
|
||||
|
||||
## コミュニティによるセッション実装
|
||||
|
||||
コミュニティは追加のセッション実装を開発しています。
|
||||
|
||||
| パッケージ | 説明 |
|
||||
|---------|-------------|
|
||||
| [openai-django-sessions](https://pypi.org/project/openai-django-sessions/) | 任意の Django 対応データベース(PostgreSQL、MySQL、SQLite など)向けの Django ORM ベースのセッション |
|
||||
|
||||
セッション実装を構築した場合は、ぜひドキュメント PR を送ってここに追加してください。
|
||||
|
||||
## API リファレンス
|
||||
|
||||
詳細な API ドキュメントについては、以下を参照してください。
|
||||
|
||||
- [`Session`][agents.memory.session.Session] - プロトコルインターフェイス
|
||||
- [`OpenAIConversationsSession`][agents.memory.OpenAIConversationsSession] - OpenAI Conversations API 実装
|
||||
- [`OpenAIResponsesCompactionSession`][agents.memory.openai_responses_compaction_session.OpenAIResponsesCompactionSession] - Responses API 圧縮ラッパー
|
||||
- [`SQLiteSession`][agents.memory.sqlite_session.SQLiteSession] - 基本的な SQLite 実装
|
||||
- [`AsyncSQLiteSession`][agents.extensions.memory.async_sqlite_session.AsyncSQLiteSession] - `aiosqlite` に基づく非同期 SQLite 実装
|
||||
- [`RedisSession`][agents.extensions.memory.redis_session.RedisSession] - Redis バックのセッション実装
|
||||
- [`SQLAlchemySession`][agents.extensions.memory.sqlalchemy_session.SQLAlchemySession] - SQLAlchemy を利用した実装
|
||||
- [`MongoDBSession`][agents.extensions.memory.mongodb_session.MongoDBSession] - MongoDB バックのセッション実装
|
||||
- [`DaprSession`][agents.extensions.memory.dapr_session.DaprSession] - Dapr ステートストア実装
|
||||
- [`AdvancedSQLiteSession`][agents.extensions.memory.advanced_sqlite_session.AdvancedSQLiteSession] - 分岐と分析を備えた拡張 SQLite
|
||||
- [`EncryptedSession`][agents.extensions.memory.encrypt_session.EncryptedSession] - 任意のセッション向けの暗号化ラッパー
|
||||
@@ -0,0 +1,80 @@
|
||||
---
|
||||
search:
|
||||
exclude: true
|
||||
---
|
||||
# SQLAlchemy セッション
|
||||
|
||||
`SQLAlchemySession` は SQLAlchemy を使用して、本番環境に対応したセッション実装を提供します。これにより、セッションストレージとして SQLAlchemy がサポートする任意のデータベース (PostgreSQL、MySQL、SQLite など) を使用できます。
|
||||
|
||||
## インストール
|
||||
|
||||
SQLAlchemy セッションには `sqlalchemy` extra が必要です:
|
||||
|
||||
```bash
|
||||
pip install openai-agents[sqlalchemy]
|
||||
```
|
||||
|
||||
## クイックスタート
|
||||
|
||||
### データベース URL の使用
|
||||
|
||||
始めるための最も簡単な方法です:
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
from agents import Agent, Runner
|
||||
from agents.extensions.memory import SQLAlchemySession
|
||||
|
||||
async def main():
|
||||
agent = Agent("Assistant")
|
||||
|
||||
# Create session using database URL
|
||||
session = SQLAlchemySession.from_url(
|
||||
"user-123",
|
||||
url="sqlite+aiosqlite:///:memory:",
|
||||
create_tables=True
|
||||
)
|
||||
|
||||
result = await Runner.run(agent, "Hello", session=session)
|
||||
print(result.final_output)
|
||||
|
||||
if __name__ == "__main__":
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
### 既存エンジンの使用
|
||||
|
||||
既存の SQLAlchemy エンジンを持つアプリケーション向けです:
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
from agents import Agent, Runner
|
||||
from agents.extensions.memory import SQLAlchemySession
|
||||
from sqlalchemy.ext.asyncio import create_async_engine
|
||||
|
||||
async def main():
|
||||
# Create your database engine
|
||||
engine = create_async_engine("postgresql+asyncpg://user:pass@localhost/db")
|
||||
|
||||
agent = Agent("Assistant")
|
||||
session = SQLAlchemySession(
|
||||
"user-456",
|
||||
engine=engine,
|
||||
create_tables=True
|
||||
)
|
||||
|
||||
result = await Runner.run(agent, "Hello", session=session)
|
||||
print(result.final_output)
|
||||
|
||||
# Clean up
|
||||
await engine.dispose()
|
||||
|
||||
if __name__ == "__main__":
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
|
||||
## API リファレンス
|
||||
|
||||
- [`SQLAlchemySession`][agents.extensions.memory.sqlalchemy_session.SQLAlchemySession] - メインクラス
|
||||
- [`Session`][agents.memory.session.Session] - 基本セッションプロトコル
|
||||
+66
-8
@@ -1,14 +1,22 @@
|
||||
---
|
||||
search:
|
||||
exclude: true
|
||||
---
|
||||
# ストリーミング
|
||||
|
||||
ストリーミングを利用すると、エージェントの実行が進行するにつれて、その更新情報を購読できます。これは、エンドユーザーに進捗状況や部分的な応答を表示するのに役立ちます。
|
||||
ストリーミングにより、エージェントの実行が進むにつれて更新を購読できます。これは、エンドユーザーに進捗状況の更新や部分的なレスポンスを表示する場合に役立ちます。
|
||||
|
||||
ストリーミングを行うには、[`Runner.run_streamed()`][agents.run.Runner.run_streamed] を呼び出します。これにより、[`RunResultStreaming`][agents.result.RunResultStreaming] が返されます。`result.stream_events()` を呼び出すと、下記で説明する [`StreamEvent`][agents.stream_events.StreamEvent] オブジェクトの非同期ストリームが得られます。
|
||||
ストリーミングするには、[`Runner.run_streamed()`][agents.run.Runner.run_streamed] を呼び出します。これにより [`RunResultStreaming`][agents.result.RunResultStreaming] が返されます。`result.stream_events()` を呼び出すと、以下で説明する [`StreamEvent`][agents.stream_events.StreamEvent] オブジェクトの非同期ストリームが得られます。
|
||||
|
||||
## raw response events
|
||||
非同期イテレーターが終了するまで、`result.stream_events()` を消費し続けてください。ストリーミング実行は、イテレーターが終了するまで完了しません。また、セッションの永続化、承認の記録管理、履歴の圧縮などの後処理は、最後の可視トークンが到着した後に完了する場合があります。ループが終了すると、`result.is_complete` は最終的な実行状態を反映します。
|
||||
|
||||
[`RawResponsesStreamEvent`][agents.stream_events.RawResponsesStreamEvent] は、LLM から直接渡される raw イベントです。これらは OpenAI Responses API フォーマットであり、各イベントには type(例:`response.created`、`response.output_text.delta` など)とデータが含まれます。これらのイベントは、応答メッセージが生成され次第、ユーザーにストリーミングしたい場合に便利です。
|
||||
## raw レスポンスイベント
|
||||
|
||||
例えば、以下の例では LLM が生成したテキストをトークンごとに出力します。
|
||||
[`RawResponsesStreamEvent`][agents.stream_events.RawResponsesStreamEvent] は、LLM から直接渡される raw イベントです。これらは OpenAI Responses API 形式であり、各イベントには型(`response.created`、`response.output_text.delta` など)とデータがあります。これらのイベントは、レスポンスメッセージが生成され次第、ユーザーにストリーミングしたい場合に役立ちます。
|
||||
|
||||
コンピュータツールの raw イベントは、保存された実行結果と同じ preview と GA の区別を維持します。Preview フローでは、1 つの `action` を持つ `computer_call` アイテムをストリーミングします。一方、`gpt-5.5` では、バッチ化された `actions[]` を持つ `computer_call` アイテムをストリーミングできます。高レベルの [`RunItemStreamEvent`][agents.stream_events.RunItemStreamEvent] サーフェスは、このために特別なコンピュータ専用イベント名を追加しません。どちらの形式も引き続き `tool_called` として表面化し、スクリーンショットの実行結果は `computer_call_output` アイテムをラップする `tool_output` として返されます。
|
||||
|
||||
たとえば、これは LLM によって生成されたテキストをトークンごとに出力します。
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
@@ -31,11 +39,61 @@ if __name__ == "__main__":
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
## Run item イベントとエージェントイベント
|
||||
## ストリーミングと承認
|
||||
|
||||
[`RunItemStreamEvent`][agents.stream_events.RunItemStreamEvent] は、より高レベルなイベントです。アイテムが完全に生成されたタイミングを通知します。これにより、「メッセージが生成された」「ツールが実行された」など、各トークン単位ではなく、進捗状況をまとめてユーザーに伝えることができます。同様に、[`AgentUpdatedStreamEvent`][agents.stream_events.AgentUpdatedStreamEvent] は、現在のエージェントが変更されたとき(例:ハンドオフの結果として)に更新情報を提供します。
|
||||
ストリーミングは、ツール承認のために一時停止する実行と互換性があります。ツールに承認が必要な場合、`result.stream_events()` は終了し、保留中の承認は [`RunResultStreaming.interruptions`][agents.result.RunResultStreaming.interruptions] で公開されます。`result.to_state()` を使って実行結果を [`RunState`][agents.run_state.RunState] に変換し、中断を承認または拒否してから、`Runner.run_streamed(...)` で再開します。
|
||||
|
||||
例えば、以下の例では raw イベントを無視し、ユーザーに更新情報のみをストリーミングします。
|
||||
```python
|
||||
result = Runner.run_streamed(agent, "Delete temporary files if they are no longer needed.")
|
||||
async for _event in result.stream_events():
|
||||
pass
|
||||
|
||||
if result.interruptions:
|
||||
state = result.to_state()
|
||||
for interruption in result.interruptions:
|
||||
state.approve(interruption)
|
||||
result = Runner.run_streamed(agent, state)
|
||||
async for _event in result.stream_events():
|
||||
pass
|
||||
```
|
||||
|
||||
一時停止/再開の完全なウォークスルーについては、[human-in-the-loop ガイド](human_in_the_loop.md)を参照してください。
|
||||
|
||||
## 現在のターン後のストリーミングのキャンセル
|
||||
|
||||
途中でストリーミング実行を停止する必要がある場合は、[`result.cancel()`][agents.result.RunResultStreaming.cancel] を呼び出します。デフォルトでは、これにより実行はすぐに停止します。停止する前に現在のターンを正常に完了させるには、代わりに `result.cancel(mode="after_turn")` を呼び出します。
|
||||
|
||||
ストリーミング実行は、`result.stream_events()` が終了するまで完了しません。最後の可視トークンの後も、SDK がセッションアイテムを永続化したり、承認状態を確定したり、履歴を圧縮したりしている場合があります。
|
||||
|
||||
[`result.to_input_list(mode="normalized")`][agents.result.RunResultBase.to_input_list] から手動で継続しており、`cancel(mode="after_turn")` がツールターンの後で停止した場合は、すぐに新しいユーザーターンを追加するのではなく、その正規化された入力で `result.last_agent` を再実行して、未完了のターンを継続してください。
|
||||
- ストリーミング実行がツール承認のために停止した場合、それを新しいターンとして扱わないでください。ストリームの読み出しを最後まで完了し、`result.interruptions` を確認して、代わりに `result.to_state()` から再開してください。
|
||||
- 次のモデル呼び出しの前に、取得したセッション履歴と新しいユーザー入力をどのようにマージするかをカスタマイズするには、[`RunConfig.session_input_callback`][agents.run.RunConfig.session_input_callback] を使用します。そこで新しいターンのアイテムを書き換えた場合、その書き換え後のバージョンがそのターンとして永続化されます。
|
||||
|
||||
## 実行アイテムイベントとエージェントイベント
|
||||
|
||||
[`RunItemStreamEvent`][agents.stream_events.RunItemStreamEvent] は、より高レベルのイベントです。アイテムが完全に生成されたタイミングを通知します。これにより、各トークン単位ではなく、「メッセージが生成された」「ツールが実行された」などのレベルで進捗更新を送信できます。同様に、[`AgentUpdatedStreamEvent`][agents.stream_events.AgentUpdatedStreamEvent] は、現在のエージェントが変更されたとき(例: ハンドオフの結果として)に更新を提供します。
|
||||
|
||||
### 実行アイテムイベント名
|
||||
|
||||
`RunItemStreamEvent.name` は、固定された一連のセマンティックなイベント名を使用します。
|
||||
|
||||
- `message_output_created`
|
||||
- `handoff_requested`
|
||||
- `handoff_occured`
|
||||
- `tool_called`
|
||||
- `tool_search_called`
|
||||
- `tool_search_output_created`
|
||||
- `tool_output`
|
||||
- `reasoning_item_created`
|
||||
- `mcp_approval_requested`
|
||||
- `mcp_approval_response`
|
||||
- `mcp_list_tools`
|
||||
|
||||
`handoff_occured` は、後方互換性のため意図的にスペルミスのままになっています。
|
||||
|
||||
ホストされたツール検索を使用する場合、モデルがツール検索リクエストを発行すると `tool_search_called` が送出され、Responses API が読み込まれたサブセットを返すと `tool_search_output_created` が送出されます。
|
||||
|
||||
たとえば、これは raw イベントを無視し、更新をユーザーにストリーミングします。
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
|
||||
+600
-35
@@ -1,18 +1,45 @@
|
||||
---
|
||||
search:
|
||||
exclude: true
|
||||
---
|
||||
# ツール
|
||||
|
||||
ツールは エージェント がアクションを実行するためのものです。たとえば、データの取得、コードの実行、外部 API の呼び出し、さらにはコンピュータ操作などが含まれます。Agents SDK には 3 種類のツールクラスがあります。
|
||||
ツールにより、エージェントはデータの取得、コードの実行、外部 API の呼び出し、さらにはコンピュータの操作といったアクションを実行できます。SDK は 5 つのカテゴリーをサポートしています。
|
||||
|
||||
- Hosted tools: これらは LLM サーバー上で AI モデルとともに実行されます。OpenAI は retrieval、Web 検索、コンピュータ操作を Hosted tools として提供しています。
|
||||
- Function calling: 任意の Python 関数をツールとして利用できます。
|
||||
- Agents as tools: エージェントをツールとして利用でき、エージェントが他のエージェントをハンドオフせずに呼び出すことができます。
|
||||
- OpenAI がホストするツール: OpenAI サーバー上でモデルと並行して実行されます。
|
||||
- ローカル/ランタイム実行ツール: `ComputerTool` と `ApplyPatchTool` は常にユーザーの環境で実行され、`ShellTool` はローカルまたはホスト型コンテナーで実行できます。
|
||||
- Function calling: 任意の Python 関数をツールとしてラップします。
|
||||
- Agents as tools: 完全なハンドオフなしで、エージェントを呼び出し可能なツールとして公開します。
|
||||
- 実験的機能: Codex ツール: ツール呼び出しから、ワークスペーススコープの Codex タスクを実行します。
|
||||
|
||||
## Hosted tools
|
||||
## ツールタイプの選択
|
||||
|
||||
OpenAI は [`OpenAIResponsesModel`][agents.models.openai_responses.OpenAIResponsesModel] を使用する際に、いくつかの組み込みツールを提供しています。
|
||||
このページをカタログとして使い、制御するランタイムに一致するセクションに移動してください。
|
||||
|
||||
- [`WebSearchTool`][agents.tool.WebSearchTool] は、エージェントが Web 検索を行うことを可能にします。
|
||||
- [`FileSearchTool`][agents.tool.FileSearchTool] は、OpenAI ベクトルストアから情報を取得できます。
|
||||
- [`ComputerTool`][agents.tool.ComputerTool] は、コンピュータ操作タスクの自動化を可能にします。
|
||||
| やりたいこと | 開始先 |
|
||||
| --- | --- |
|
||||
| OpenAI 管理のツール(Web 検索、ファイル検索、code interpreter、ホスト型 MCP、画像生成)を使用する | [ホスト型ツール](#hosted-tools) |
|
||||
| ツール検索で大規模なツールサーフェスをランタイムまで遅延させる | [ホスト型ツール検索](#hosted-tool-search) |
|
||||
| 自身のプロセスまたは環境でツールを実行する | [ローカルランタイムツール](#local-runtime-tools) |
|
||||
| Python 関数をツールとしてラップする | [関数ツール](#function-tools) |
|
||||
| ハンドオフなしで、あるエージェントが別のエージェントを呼び出せるようにする | [Agents as tools](#agents-as-tools) |
|
||||
| エージェントからワークスペーススコープの Codex タスクを実行する | [実験的機能: Codex ツール](#experimental-codex-tool) |
|
||||
|
||||
## ホスト型ツール
|
||||
|
||||
OpenAI は [`OpenAIResponsesModel`][agents.models.openai_responses.OpenAIResponsesModel] を使用する場合に、いくつかの組み込みツールを提供しています。
|
||||
|
||||
- [`WebSearchTool`][agents.tool.WebSearchTool] により、エージェントは Web を検索できます。
|
||||
- [`FileSearchTool`][agents.tool.FileSearchTool] により、OpenAI ベクトルストアから情報を取得できます。
|
||||
- [`CodeInterpreterTool`][agents.tool.CodeInterpreterTool] により、LLM はサンドボックス化された環境でコードを実行できます。
|
||||
- [`HostedMCPTool`][agents.tool.HostedMCPTool] は、リモート MCP サーバーのツールをモデルに公開します。
|
||||
- [`ImageGenerationTool`][agents.tool.ImageGenerationTool] は、プロンプトから画像を生成します。
|
||||
- [`ToolSearchTool`][agents.tool.ToolSearchTool] により、モデルは遅延されたツール、名前空間、またはホスト型 MCP サーバーを必要に応じて読み込めます。
|
||||
|
||||
高度なホスト型検索オプション:
|
||||
|
||||
- `FileSearchTool` は、`vector_store_ids` と `max_num_results` に加えて、`filters`、`ranking_options`、`include_search_results` をサポートします。
|
||||
- `WebSearchTool` は、`filters`、`user_location`、`search_context_size` をサポートします。
|
||||
|
||||
```python
|
||||
from agents import Agent, FileSearchTool, Runner, WebSearchTool
|
||||
@@ -33,16 +60,203 @@ async def main():
|
||||
print(result.final_output)
|
||||
```
|
||||
|
||||
## Function tools
|
||||
### ホスト型ツール検索
|
||||
|
||||
任意の Python 関数をツールとして利用できます。Agents SDK が自動的にツールをセットアップします。
|
||||
ツール検索により、OpenAI Responses モデルは大規模なツールサーフェスをランタイムまで遅延できるため、モデルは現在のターンに必要なサブセットのみを読み込みます。これは、多数の関数ツール、名前空間グループ、またはホスト型 MCP サーバーがあり、すべてのツールを事前に公開せずにツールスキーマのトークンを削減したい場合に便利です。
|
||||
|
||||
- ツール名は Python 関数名になります(または任意の名前を指定できます)
|
||||
- ツールの説明は関数の docstring から取得されます(または任意の説明を指定できます)
|
||||
- 関数の引数から自動的に入力スキーマが作成されます
|
||||
- 各入力の説明は、関数の docstring から取得されます(無効化も可能です)
|
||||
エージェントを構築する時点で候補ツールがすでに分かっている場合は、ホスト型ツール検索から始めてください。アプリケーションが何を読み込むかを動的に決定する必要がある場合、Responses API はクライアント実行型ツール検索もサポートしていますが、標準の `Runner` はそのモードを自動実行しません。
|
||||
|
||||
Python の `inspect` モジュールを使って関数シグネチャを抽出し、[`griffe`](https://mkdocstrings.github.io/griffe/) で docstring を解析し、`pydantic` でスキーマを作成します。
|
||||
```python
|
||||
from typing import Annotated
|
||||
|
||||
from agents import Agent, Runner, ToolSearchTool, function_tool, tool_namespace
|
||||
|
||||
|
||||
@function_tool(defer_loading=True)
|
||||
def get_customer_profile(
|
||||
customer_id: Annotated[str, "The customer ID to look up."],
|
||||
) -> str:
|
||||
"""Fetch a CRM customer profile."""
|
||||
return f"profile for {customer_id}"
|
||||
|
||||
|
||||
@function_tool(defer_loading=True)
|
||||
def list_open_orders(
|
||||
customer_id: Annotated[str, "The customer ID to look up."],
|
||||
) -> str:
|
||||
"""List open orders for a customer."""
|
||||
return f"open orders for {customer_id}"
|
||||
|
||||
|
||||
crm_tools = tool_namespace(
|
||||
name="crm",
|
||||
description="CRM tools for customer lookups.",
|
||||
tools=[get_customer_profile, list_open_orders],
|
||||
)
|
||||
|
||||
|
||||
agent = Agent(
|
||||
name="Operations assistant",
|
||||
model="gpt-5.5",
|
||||
instructions="Load the crm namespace before using CRM tools.",
|
||||
tools=[*crm_tools, ToolSearchTool()],
|
||||
)
|
||||
|
||||
result = await Runner.run(agent, "Look up customer_42 and list their open orders.")
|
||||
print(result.final_output)
|
||||
```
|
||||
|
||||
知っておくべきこと:
|
||||
|
||||
- ホスト型ツール検索は、OpenAI Responses モデルでのみ利用できます。現在の Python SDK のサポートは `openai>=2.25.0` に依存します。
|
||||
- エージェントで遅延読み込みサーフェスを構成するときは、`ToolSearchTool()` を正確に 1 つ追加してください。
|
||||
- 検索可能なサーフェスには、`@function_tool(defer_loading=True)`、`tool_namespace(name=..., description=..., tools=[...])`、`HostedMCPTool(tool_config={..., "defer_loading": True})` が含まれます。
|
||||
- 遅延読み込みの関数ツールは、`ToolSearchTool()` と組み合わせる必要があります。名前空間のみの構成でも、モデルが必要に応じて適切なグループを読み込めるようにするために `ToolSearchTool()` を使用できます。
|
||||
- `tool_namespace()` は、`FunctionTool` インスタンスを共有の名前空間名と説明の下にグループ化します。これは通常、`crm`、`billing`、`shipping` など、関連するツールが多数ある場合に最適です。
|
||||
- OpenAI の公式ベストプラクティスガイダンスは、[可能な場合は名前空間を使用する](https://developers.openai.com/api/docs/guides/tools-tool-search#use-namespaces-where-possible)です。
|
||||
- 可能な場合は、多数の個別に遅延された関数よりも、名前空間またはホスト型 MCP サーバーを優先してください。通常、これらはモデルに対してより優れた高レベルの検索サーフェスと、より大きなトークン節約を提供します。
|
||||
- 名前空間では、即時ツールと遅延ツールを混在させることができます。`defer_loading=True` のないツールはすぐに呼び出し可能なままで、同じ名前空間内の遅延ツールはツール検索を通じて読み込まれます。
|
||||
- 目安として、各名前空間はかなり小さく保ち、理想的には 10 個未満の関数にしてください。
|
||||
- 名前付きの `tool_choice` は、裸の名前空間名や遅延のみのツールを対象にできません。`auto`、`required`、または実際のトップレベルの呼び出し可能なツール名を優先してください。
|
||||
- `ToolSearchTool(execution="client")` は、手動の Responses オーケストレーション用です。モデルがクライアント実行型の `tool_search_call` を発行した場合、標準の `Runner` はそれを実行せずに例外を送出します。
|
||||
- ツール検索アクティビティは、[`RunResult.new_items`](results.md#new-items) と [`RunItemStreamEvent`](streaming.md#run-item-event-names) に、専用の項目タイプとイベントタイプで表示されます。
|
||||
- 名前空間付き読み込みとトップレベルの遅延ツールの両方を扱う、完全に実行可能なコード例については、`examples/tools/tool_search.py` を参照してください。
|
||||
- 公式プラットフォームガイド: [ツール検索](https://developers.openai.com/api/docs/guides/tools-tool-search)。
|
||||
|
||||
### ホスト型コンテナーシェル + スキル
|
||||
|
||||
`ShellTool` は、OpenAI がホストするコンテナー実行もサポートしています。ローカルランタイムではなく、管理されたコンテナー内でモデルにシェルコマンドを実行させたい場合は、このモードを使用します。
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner, ShellTool, ShellToolSkillReference
|
||||
|
||||
csv_skill: ShellToolSkillReference = {
|
||||
"type": "skill_reference",
|
||||
"skill_id": "skill_698bbe879adc81918725cbc69dcae7960bc5613dadaed377",
|
||||
"version": "1",
|
||||
}
|
||||
|
||||
agent = Agent(
|
||||
name="Container shell agent",
|
||||
model="gpt-5.5",
|
||||
instructions="Use the mounted skill when helpful.",
|
||||
tools=[
|
||||
ShellTool(
|
||||
environment={
|
||||
"type": "container_auto",
|
||||
"network_policy": {"type": "disabled"},
|
||||
"skills": [csv_skill],
|
||||
}
|
||||
)
|
||||
],
|
||||
)
|
||||
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
"Use the configured skill to analyze CSV files in /mnt/data and summarize totals by region.",
|
||||
)
|
||||
print(result.final_output)
|
||||
```
|
||||
|
||||
以降の実行で既存のコンテナーを再利用するには、`environment={"type": "container_reference", "container_id": "cntr_..."}` を設定します。
|
||||
|
||||
知っておくべきこと:
|
||||
|
||||
- ホスト型シェルは、Responses API シェルツールを通じて利用できます。
|
||||
- `container_auto` は、リクエスト用のコンテナーをプロビジョニングします。`container_reference` は既存のコンテナーを再利用します。
|
||||
- `container_auto` には `file_ids` と `memory_limit` も含めることができます。
|
||||
- `environment.skills` は、スキル参照とインラインスキルバンドルを受け取ります。
|
||||
- ホスト型環境では、`ShellTool` に `executor`、`needs_approval`、`on_approval` を設定しないでください。
|
||||
- `network_policy` は、`disabled` モードと `allowlist` モードをサポートします。
|
||||
- `allowlist` モードでは、`network_policy.domain_secrets` により、名前でドメインスコープのシークレットを注入できます。
|
||||
- 完全なコード例については、`examples/tools/container_shell_skill_reference.py` と `examples/tools/container_shell_inline_skill.py` を参照してください。
|
||||
- OpenAI プラットフォームガイド: [シェル](https://platform.openai.com/docs/guides/tools-shell) と [スキル](https://platform.openai.com/docs/guides/tools-skills)。
|
||||
|
||||
## ローカルランタイムツール
|
||||
|
||||
ローカルランタイムツールは、モデル応答自体の外で実行されます。モデルは引き続きいつ呼び出すかを決定しますが、実際の作業はアプリケーションまたは構成された実行環境が実行します。
|
||||
|
||||
`ComputerTool` と `ApplyPatchTool` には、常にユーザーが提供するローカル実装が必要です。`ShellTool` は両方のモードにまたがっています。管理された実行が必要な場合は上記のホスト型コンテナー構成を使用し、自身のプロセスでコマンドを実行したい場合は下記のローカルランタイム構成を使用してください。
|
||||
|
||||
ローカルランタイムツールでは、実装を提供する必要があります。
|
||||
|
||||
- [`ComputerTool`][agents.tool.ComputerTool]: GUI/ブラウザー自動化を有効にするために、[`Computer`][agents.computer.Computer] または [`AsyncComputer`][agents.computer.AsyncComputer] インターフェイスを実装します。
|
||||
- [`ShellTool`][agents.tool.ShellTool]: ローカル実行とホスト型コンテナー実行の両方に対応する最新のシェルツールです。
|
||||
- [`LocalShellTool`][agents.tool.LocalShellTool]: レガシーなローカルシェル連携です。
|
||||
- [`ApplyPatchTool`][agents.tool.ApplyPatchTool]: diff をローカルに適用するために [`ApplyPatchEditor`][agents.editor.ApplyPatchEditor] を実装します。
|
||||
- ローカルシェルスキルは、`ShellTool(environment={"type": "local", "skills": [...]})` で利用できます。
|
||||
|
||||
### ComputerTool と Responses コンピュータツール
|
||||
|
||||
`ComputerTool` は引き続きローカルハーネスです。[`Computer`][agents.computer.Computer] または [`AsyncComputer`][agents.computer.AsyncComputer] の実装を提供すると、SDK はそのハーネスを OpenAI Responses API のコンピュータサーフェスにマッピングします。
|
||||
|
||||
明示的な [`gpt-5.5`](https://developers.openai.com/api/docs/models/gpt-5.5) リクエストでは、SDK は GA 組み込みツールペイロード `{"type": "computer"}` を送信します。古い `computer-use-preview` モデルでは、プレビューペイロード `{"type": "computer_use_preview", "environment": ..., "display_width": ..., "display_height": ...}` が維持されます。これは、OpenAI の [コンピュータ操作ガイド](https://developers.openai.com/api/docs/guides/tools-computer-use/)で説明されているプラットフォーム移行を反映しています。
|
||||
|
||||
- モデル: `computer-use-preview` -> `gpt-5.5`
|
||||
- ツールセレクター: `computer_use_preview` -> `computer`
|
||||
- コンピュータ呼び出しの形状: `computer_call` ごとに 1 つの `action` -> `computer_call` 上のバッチ化された `actions[]`
|
||||
- 切り捨て: プレビューパスでは `ModelSettings(truncation="auto")` が必須 -> GA パスでは不要
|
||||
|
||||
SDK は、実際の Responses リクエスト上の有効なモデルから、そのワイヤ形式を選択します。プロンプトテンプレートを使用し、プロンプトがモデルを保持しているためリクエストが `model` を省略する場合、`model="gpt-5.5"` を明示的に保持するか、`ModelSettings(tool_choice="computer")` または `ModelSettings(tool_choice="computer_use")` で GA セレクターを強制しない限り、SDK はプレビュー互換のコンピュータペイロードを維持します。
|
||||
|
||||
[`ComputerTool`][agents.tool.ComputerTool] が存在する場合、`tool_choice="computer"`、`"computer_use"`、`"computer_use_preview"` はすべて受け入れられ、有効なリクエストモデルに一致する組み込みセレクターに正規化されます。`ComputerTool` がない場合、これらの文字列は引き続き通常の関数名のように動作します。
|
||||
|
||||
この違いは、`ComputerTool` が [`ComputerProvider`][agents.tool.ComputerProvider] ファクトリーによって支えられている場合に重要です。GA の `computer` ペイロードはシリアライズ時に `environment` や寸法を必要としないため、未解決のファクトリーでも問題ありません。プレビュー互換のシリアライズでは、SDK が `environment`、`display_width`、`display_height` を送信できるように、解決済みの `Computer` または `AsyncComputer` インスタンスが引き続き必要です。
|
||||
|
||||
ランタイムでは、どちらのパスも同じローカルハーネスを使用します。プレビュー応答は単一の `action` を持つ `computer_call` 項目を発行します。`gpt-5.5` はバッチ化された `actions[]` を発行でき、SDK は `computer_call_output` スクリーンショット項目を生成する前にそれらを順番に実行します。実行可能な Playwright ベースのハーネスについては、`examples/tools/computer_use.py` を参照してください。
|
||||
|
||||
```python
|
||||
from agents import Agent, ApplyPatchTool, ShellTool
|
||||
from agents.computer import AsyncComputer
|
||||
from agents.editor import ApplyPatchResult, ApplyPatchOperation, ApplyPatchEditor
|
||||
|
||||
|
||||
class NoopComputer(AsyncComputer):
|
||||
environment = "browser"
|
||||
dimensions = (1024, 768)
|
||||
async def screenshot(self): return ""
|
||||
async def click(self, x, y, button): ...
|
||||
async def double_click(self, x, y): ...
|
||||
async def scroll(self, x, y, scroll_x, scroll_y): ...
|
||||
async def type(self, text): ...
|
||||
async def wait(self): ...
|
||||
async def move(self, x, y): ...
|
||||
async def keypress(self, keys): ...
|
||||
async def drag(self, path): ...
|
||||
|
||||
|
||||
class NoopEditor(ApplyPatchEditor):
|
||||
async def create_file(self, op: ApplyPatchOperation): return ApplyPatchResult(status="completed")
|
||||
async def update_file(self, op: ApplyPatchOperation): return ApplyPatchResult(status="completed")
|
||||
async def delete_file(self, op: ApplyPatchOperation): return ApplyPatchResult(status="completed")
|
||||
|
||||
|
||||
async def run_shell(request):
|
||||
return "shell output"
|
||||
|
||||
|
||||
agent = Agent(
|
||||
name="Local tools agent",
|
||||
tools=[
|
||||
ShellTool(executor=run_shell),
|
||||
ApplyPatchTool(editor=NoopEditor()),
|
||||
# ComputerTool expects a Computer/AsyncComputer implementation; omitted here for brevity.
|
||||
],
|
||||
)
|
||||
```
|
||||
|
||||
## 関数ツール
|
||||
|
||||
任意の Python 関数をツールとして使用できます。Agents SDK はツールを自動的に設定します。
|
||||
|
||||
- ツールの名前は Python 関数の名前になります(または名前を指定できます)
|
||||
- ツールの説明は関数の docstring から取得されます(または説明を指定できます)
|
||||
- 関数入力のスキーマは、関数の引数から自動的に作成されます
|
||||
- 各入力の説明は、無効化されていない限り、関数の docstring から取得されます
|
||||
|
||||
関数シグネチャの抽出には Python の `inspect` モジュールを使用し、docstring の解析には [`griffe`](https://mkdocstrings.github.io/griffe/) を、スキーマ作成には `pydantic` を使用します。
|
||||
|
||||
OpenAI Responses モデルを使用している場合、`@function_tool(defer_loading=True)` は `ToolSearchTool()` が読み込むまで関数ツールを隠します。[`tool_namespace()`][agents.tool.tool_namespace] を使用して、関連する関数ツールをグループ化することもできます。詳細な設定と制約については、[ホスト型ツール検索](#hosted-tool-search)を参照してください。
|
||||
|
||||
```python
|
||||
import json
|
||||
@@ -94,12 +308,12 @@ for tool in agent.tools:
|
||||
|
||||
```
|
||||
|
||||
1. 関数の引数には任意の Python 型を利用でき、同期・非同期どちらの関数も利用可能です。
|
||||
2. docstring があれば、説明や引数の説明として利用されます。
|
||||
3. 関数はオプションで `context` を最初の引数として受け取れます。また、ツール名や説明、docstring スタイルの指定などのオーバーライドも可能です。
|
||||
4. デコレートした関数をツールのリストに渡すことができます。
|
||||
1. 関数の引数には任意の Python 型を使用でき、関数は同期でも非同期でもかまいません。
|
||||
2. Docstring が存在する場合、説明と引数の説明を取得するために使用されます
|
||||
3. 関数は任意で `context` を受け取れます(最初の引数である必要があります)。ツール名、説明、使用する docstring スタイルなどのオーバーライドも設定できます。
|
||||
4. デコレートされた関数をツールのリストに渡すことができます。
|
||||
|
||||
??? note "出力を展開して表示"
|
||||
??? note "出力を表示するには展開してください"
|
||||
|
||||
```
|
||||
fetch_weather
|
||||
@@ -169,14 +383,22 @@ for tool in agent.tools:
|
||||
}
|
||||
```
|
||||
|
||||
### カスタム function tools
|
||||
### 関数ツールからの画像またはファイルの返却
|
||||
|
||||
場合によっては、Python 関数をツールとして使いたくないこともあります。その場合は、[`FunctionTool`][agents.tool.FunctionTool] を直接作成できます。必要な情報は以下の通りです。
|
||||
テキスト出力を返すことに加えて、関数ツールの出力として 1 つまたは複数の画像やファイルを返すことができます。そのためには、次のいずれかを返せます。
|
||||
|
||||
- 画像: [`ToolOutputImage`][agents.tool.ToolOutputImage](または TypedDict 版の [`ToolOutputImageDict`][agents.tool.ToolOutputImageDict])
|
||||
- ファイル: [`ToolOutputFileContent`][agents.tool.ToolOutputFileContent](または TypedDict 版の [`ToolOutputFileContentDict`][agents.tool.ToolOutputFileContentDict])
|
||||
- テキスト: 文字列または文字列化可能なオブジェクト、または [`ToolOutputText`][agents.tool.ToolOutputText](または TypedDict 版の [`ToolOutputTextDict`][agents.tool.ToolOutputTextDict])
|
||||
|
||||
### カスタム関数ツール
|
||||
|
||||
Python 関数をツールとして使いたくない場合があります。希望する場合は、[`FunctionTool`][agents.tool.FunctionTool] を直接作成できます。次のものを提供する必要があります。
|
||||
|
||||
- `name`
|
||||
- `description`
|
||||
- `params_json_schema`(引数の JSON スキーマ)
|
||||
- `on_invoke_tool`(context と引数(JSON 文字列)を受け取り、ツールの出力を文字列で返す async 関数)
|
||||
- `params_json_schema`: 引数用の JSON スキーマです
|
||||
- `on_invoke_tool`: [`ToolContext`][agents.tool_context.ToolContext] と引数を JSON 文字列として受け取り、ツール出力(たとえば、テキスト、構造化ツール出力オブジェクト、または出力のリスト)を返す非同期関数です。
|
||||
|
||||
```python
|
||||
from typing import Any
|
||||
@@ -211,16 +433,120 @@ tool = FunctionTool(
|
||||
|
||||
### 引数と docstring の自動解析
|
||||
|
||||
前述の通り、関数シグネチャを自動解析してツールのスキーマを抽出し、docstring からツールや各引数の説明を抽出します。主なポイントは以下の通りです。
|
||||
前述のとおり、ツールのスキーマを抽出するために関数シグネチャを自動的に解析し、ツールと個々の引数の説明を抽出するために docstring を解析します。これについての注意点は次のとおりです。
|
||||
|
||||
1. シグネチャの解析は `inspect` モジュールで行います。型アノテーションを利用して引数の型を把握し、Pydantic モデルを動的に構築して全体のスキーマを表現します。Python の基本コンポーネント、Pydantic モデル、TypedDict など、ほとんどの型をサポートしています。
|
||||
2. docstring の解析には `griffe` を使用します。サポートされている docstring フォーマットは `google`、`sphinx`、`numpy` です。docstring フォーマットは自動検出を試みますが、`function_tool` 呼び出し時に明示的に指定することもできます。`use_docstring_info` を `False` に設定することで docstring 解析を無効化できます。
|
||||
1. シグネチャ解析は `inspect` モジュールを介して行われます。引数の型を理解するために型アノテーションを使用し、全体のスキーマを表す Pydantic モデルを動的に構築します。Python のプリミティブ型、Pydantic モデル、TypedDict など、ほとんどの型をサポートします。
|
||||
2. Docstring の解析には `griffe` を使用します。サポートされる docstring 形式は `google`、`sphinx`、`numpy` です。docstring 形式の自動検出を試みますが、これはベストエフォートであり、`function_tool` を呼び出すときに明示的に設定できます。`use_docstring_info` を `False` に設定することで、docstring 解析を無効化することもできます。
|
||||
|
||||
スキーマ抽出のコードは [`agents.function_schema`][] にあります。
|
||||
|
||||
### Pydantic Field による引数の制約と説明
|
||||
|
||||
Pydantic の [`Field`](https://docs.pydantic.dev/latest/concepts/fields/) を使用して、ツール引数に制約(例: 数値の最小/最大、文字列の長さまたはパターン)と説明を追加できます。Pydantic と同様に、デフォルトベース(`arg: int = Field(..., ge=1)`)と `Annotated`(`arg: Annotated[int, Field(..., ge=1)]`)の両方の形式がサポートされます。生成される JSON スキーマと検証には、これらの制約が含まれます。
|
||||
|
||||
```python
|
||||
from typing import Annotated
|
||||
from pydantic import Field
|
||||
from agents import function_tool
|
||||
|
||||
# Default-based form
|
||||
@function_tool
|
||||
def score_a(score: int = Field(..., ge=0, le=100, description="Score from 0 to 100")) -> str:
|
||||
return f"Score recorded: {score}"
|
||||
|
||||
# Annotated form
|
||||
@function_tool
|
||||
def score_b(score: Annotated[int, Field(..., ge=0, le=100, description="Score from 0 to 100")]) -> str:
|
||||
return f"Score recorded: {score}"
|
||||
```
|
||||
|
||||
### 関数ツールのタイムアウト
|
||||
|
||||
`@function_tool(timeout=...)` を使用して、非同期関数ツールに呼び出し単位のタイムアウトを設定できます。
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
from agents import Agent, Runner, function_tool
|
||||
|
||||
|
||||
@function_tool(timeout=2.0)
|
||||
async def slow_lookup(query: str) -> str:
|
||||
await asyncio.sleep(10)
|
||||
return f"Result for {query}"
|
||||
|
||||
|
||||
agent = Agent(
|
||||
name="Timeout demo",
|
||||
instructions="Use tools when helpful.",
|
||||
tools=[slow_lookup],
|
||||
)
|
||||
```
|
||||
|
||||
タイムアウトに達した場合、デフォルトの動作は `timeout_behavior="error_as_result"` で、モデルに表示されるタイムアウトメッセージ(例: `Tool 'slow_lookup' timed out after 2 seconds.`)を送信します。
|
||||
|
||||
タイムアウト処理を制御できます。
|
||||
|
||||
- `timeout_behavior="error_as_result"`(デフォルト): モデルが回復できるように、タイムアウトメッセージをモデルに返します。
|
||||
- `timeout_behavior="raise_exception"`: [`ToolTimeoutError`][agents.exceptions.ToolTimeoutError] を送出し、実行を失敗させます。
|
||||
- `timeout_error_function=...`: `error_as_result` を使用する場合のタイムアウトメッセージをカスタマイズします。
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
from agents import Agent, Runner, ToolTimeoutError, function_tool
|
||||
|
||||
|
||||
@function_tool(timeout=1.5, timeout_behavior="raise_exception")
|
||||
async def slow_tool() -> str:
|
||||
await asyncio.sleep(5)
|
||||
return "done"
|
||||
|
||||
|
||||
agent = Agent(name="Timeout hard-fail", tools=[slow_tool])
|
||||
|
||||
try:
|
||||
await Runner.run(agent, "Run the tool")
|
||||
except ToolTimeoutError as e:
|
||||
print(f"{e.tool_name} timed out in {e.timeout_seconds} seconds")
|
||||
```
|
||||
|
||||
!!! note
|
||||
|
||||
タイムアウト構成は、非同期の `@function_tool` ハンドラーでのみサポートされます。
|
||||
|
||||
### 関数ツールのエラー処理
|
||||
|
||||
`@function_tool` を介して関数ツールを作成する場合、`failure_error_function` を渡すことができます。これは、ツール呼び出しがクラッシュした場合に LLM へエラー応答を提供する関数です。
|
||||
|
||||
- デフォルトでは(つまり何も渡さない場合)、LLM にエラーが発生したことを伝える `default_tool_error_function` が実行されます。
|
||||
- 独自のエラー関数を渡した場合は、それが代わりに実行され、その応答が LLM に送信されます。
|
||||
- 明示的に `None` を渡した場合、任意のツール呼び出しエラーが再送出され、呼び出し側で処理できます。これは、モデルが無効な JSON を生成した場合の `ModelBehaviorError` や、コードがクラッシュした場合の `UserError` などになり得ます。
|
||||
|
||||
```python
|
||||
from agents import function_tool, RunContextWrapper
|
||||
from typing import Any
|
||||
|
||||
def my_custom_error_function(context: RunContextWrapper[Any], error: Exception) -> str:
|
||||
"""A custom function to provide a user-friendly error message."""
|
||||
print(f"A tool call failed with the following error: {error}")
|
||||
return "An internal server error occurred. Please try again later."
|
||||
|
||||
@function_tool(failure_error_function=my_custom_error_function)
|
||||
def get_user_profile(user_id: str) -> str:
|
||||
"""Fetches a user profile from a mock API.
|
||||
This function demonstrates a 'flaky' or failing API call.
|
||||
"""
|
||||
if user_id == "user_123":
|
||||
return "User profile for user_123 successfully retrieved."
|
||||
else:
|
||||
raise ValueError(f"Could not retrieve profile for user_id: {user_id}. API returned an error.")
|
||||
|
||||
```
|
||||
|
||||
`FunctionTool` オブジェクトを手動で作成している場合は、`on_invoke_tool` 関数内でエラーを処理する必要があります。
|
||||
|
||||
## Agents as tools
|
||||
|
||||
一部のワークフローでは、ハンドオフせずに中央のエージェントが専門エージェントのネットワークをオーケストレーションしたい場合があります。その場合、エージェントをツールとしてモデル化することで実現できます。
|
||||
一部のワークフローでは、制御をハンドオフする代わりに、中央エージェントに専門特化したエージェントのネットワークをオーケストレーションさせたい場合があります。これは、エージェントをツールとしてモデル化することで実現できます。
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner
|
||||
@@ -259,12 +585,251 @@ async def main():
|
||||
print(result.final_output)
|
||||
```
|
||||
|
||||
## Function tools でのエラー処理
|
||||
### ツールエージェントのカスタマイズ
|
||||
|
||||
`@function_tool` で function tool を作成する際、`failure_error_function` を渡すことができます。これは、ツール呼び出しがクラッシュした場合に LLM へエラー応答を提供する関数です。
|
||||
`agent.as_tool` 関数は、エージェントをツールに簡単に変換するための便利なメソッドです。`max_turns`、`run_config`、`hooks`、`previous_response_id`、`conversation_id`、`session`、`needs_approval` などの一般的なランタイムオプションをサポートします。また、`parameters`、`input_builder`、`include_input_schema` による構造化入力もサポートします。高度なオーケストレーション(たとえば、条件付きリトライ、フォールバック動作、複数のエージェント呼び出しのチェーン)の場合は、ツール実装内で `Runner.run` を直接使用してください。
|
||||
|
||||
- デフォルト(何も渡さない場合)では、`default_tool_error_function` が実行され、LLM にエラーが発生したことを伝えます。
|
||||
- 独自のエラー関数を渡した場合は、それが実行され、その応答が LLM に送信されます。
|
||||
- 明示的に `None` を渡した場合、ツール呼び出し時のエラーは再スローされ、ユーザー側で処理できます。たとえば、モデルが無効な JSON を生成した場合は `ModelBehaviorError`、コードがクラッシュした場合は `UserError` などです。
|
||||
```python
|
||||
@function_tool
|
||||
async def run_my_agent() -> str:
|
||||
"""A tool that runs the agent with custom configs"""
|
||||
|
||||
`FunctionTool` オブジェクトを手動で作成する場合は、`on_invoke_tool` 関数内でエラー処理を行う必要があります。
|
||||
agent = Agent(name="My agent", instructions="...")
|
||||
|
||||
result = await Runner.run(
|
||||
agent,
|
||||
input="...",
|
||||
max_turns=5,
|
||||
run_config=...
|
||||
)
|
||||
|
||||
return str(result.final_output)
|
||||
```
|
||||
|
||||
### ツールエージェントの構造化入力
|
||||
|
||||
デフォルトでは、`Agent.as_tool()` は単一の文字列入力(`{"input": "..."}`)を想定しますが、`parameters`(Pydantic モデルまたは dataclass 型)を渡すことで構造化スキーマを公開できます。
|
||||
|
||||
追加オプション:
|
||||
|
||||
- `include_input_schema=True` は、生成されるネストされた入力に完全な JSON スキーマを含めます。
|
||||
- `input_builder=...` により、構造化ツール引数がネストされたエージェント入力になる方法を完全にカスタマイズできます。
|
||||
- `RunContextWrapper.tool_input` には、ネストされた実行コンテキスト内の解析済み構造化ペイロードが含まれます。
|
||||
|
||||
```python
|
||||
from pydantic import BaseModel, Field
|
||||
|
||||
|
||||
class TranslationInput(BaseModel):
|
||||
text: str = Field(description="Text to translate.")
|
||||
source: str = Field(description="Source language.")
|
||||
target: str = Field(description="Target language.")
|
||||
|
||||
|
||||
translator_tool = translator_agent.as_tool(
|
||||
tool_name="translate_text",
|
||||
tool_description="Translate text between languages.",
|
||||
parameters=TranslationInput,
|
||||
include_input_schema=True,
|
||||
)
|
||||
```
|
||||
|
||||
完全に実行可能な例については、`examples/agent_patterns/agents_as_tools_structured.py` を参照してください。
|
||||
|
||||
### ツールエージェントの承認ゲート
|
||||
|
||||
`Agent.as_tool(..., needs_approval=...)` は、`function_tool` と同じ承認フローを使用します。承認が必要な場合、実行は一時停止し、保留中の項目が `result.interruptions` に表示されます。その後、`result.to_state()` を使い、`state.approve(...)` または `state.reject(...)` を呼び出した後に再開します。完全な一時停止/再開パターンについては、[Human-in-the-loop ガイド](human_in_the_loop.md)を参照してください。
|
||||
|
||||
### カスタム出力抽出
|
||||
|
||||
場合によっては、中央エージェントに返す前に、ツールエージェントの出力を変更したいことがあります。これは、次のような場合に便利です。
|
||||
|
||||
- サブエージェントのチャット履歴から特定の情報(例: JSON ペイロード)を抽出する。
|
||||
- エージェントの最終回答を変換または再フォーマットする(例: Markdown をプレーンテキストまたは CSV に変換する)。
|
||||
- エージェントの応答が欠落している、または不正な形式の場合に、出力を検証するかフォールバック値を提供する。
|
||||
|
||||
これは、`as_tool` メソッドに `custom_output_extractor` 引数を指定することで実行できます。
|
||||
|
||||
```python
|
||||
async def extract_json_payload(run_result: RunResult) -> str:
|
||||
# Scan the agent’s outputs in reverse order until we find a JSON-like message from a tool call.
|
||||
for item in reversed(run_result.new_items):
|
||||
if isinstance(item, ToolCallOutputItem) and item.output.strip().startswith("{"):
|
||||
return item.output.strip()
|
||||
# Fallback to an empty JSON object if nothing was found
|
||||
return "{}"
|
||||
|
||||
|
||||
json_tool = data_agent.as_tool(
|
||||
tool_name="get_data_json",
|
||||
tool_description="Run the data agent and return only its JSON payload",
|
||||
custom_output_extractor=extract_json_payload,
|
||||
)
|
||||
```
|
||||
|
||||
カスタム抽出器内では、ネストされた [`RunResult`][agents.result.RunResult] も
|
||||
[`agent_tool_invocation`][agents.result.RunResultBase.agent_tool_invocation] を公開します。これは、ネストされた実行結果を後処理する際に、
|
||||
外側のツール名、呼び出し ID、または生の引数が必要な場合に便利です。
|
||||
[実行結果ガイド](results.md#agent-as-tool-metadata)を参照してください。
|
||||
|
||||
### ネストされたエージェント実行のストリーミング
|
||||
|
||||
`as_tool` に `on_stream` コールバックを渡すと、ネストされたエージェントが発行するストリーミングイベントをリッスンしつつ、ストリーム完了後にその最終出力を返せます。
|
||||
|
||||
```python
|
||||
from agents import AgentToolStreamEvent
|
||||
|
||||
|
||||
async def handle_stream(event: AgentToolStreamEvent) -> None:
|
||||
# Inspect the underlying StreamEvent along with agent metadata.
|
||||
print(f"[stream] {event['agent'].name} :: {event['event'].type}")
|
||||
|
||||
|
||||
billing_agent_tool = billing_agent.as_tool(
|
||||
tool_name="billing_helper",
|
||||
tool_description="Answer billing questions.",
|
||||
on_stream=handle_stream, # Can be sync or async.
|
||||
)
|
||||
```
|
||||
|
||||
想定されること:
|
||||
|
||||
- イベントタイプは `StreamEvent["type"]` を反映します: `raw_response_event`、`run_item_stream_event`、`agent_updated_stream_event`。
|
||||
- `on_stream` を指定すると、ネストされたエージェントは自動的にストリーミングモードで実行され、最終出力を返す前にストリームが読み切られます。
|
||||
- ハンドラーは同期または非同期にできます。各イベントは到着した順に配信されます。
|
||||
- モデルツール呼び出しを介してツールが呼び出された場合、`tool_call` が存在します。直接呼び出しでは `None` のままになる場合があります。
|
||||
- 完全に実行可能なサンプルについては、`examples/agent_patterns/agents_as_tools_streaming.py` を参照してください。
|
||||
|
||||
### 条件付きツール有効化
|
||||
|
||||
`is_enabled` パラメーターを使用して、ランタイムでエージェントツールを条件付きで有効化または無効化できます。これにより、コンテキスト、ユーザー設定、またはランタイム条件に基づいて、LLM が利用できるツールを動的にフィルタリングできます。
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
from agents import Agent, AgentBase, Runner, RunContextWrapper
|
||||
from pydantic import BaseModel
|
||||
|
||||
class LanguageContext(BaseModel):
|
||||
language_preference: str = "french_spanish"
|
||||
|
||||
def french_enabled(ctx: RunContextWrapper[LanguageContext], agent: AgentBase) -> bool:
|
||||
"""Enable French for French+Spanish preference."""
|
||||
return ctx.context.language_preference == "french_spanish"
|
||||
|
||||
# Create specialized agents
|
||||
spanish_agent = Agent(
|
||||
name="spanish_agent",
|
||||
instructions="You respond in Spanish. Always reply to the user's question in Spanish.",
|
||||
)
|
||||
|
||||
french_agent = Agent(
|
||||
name="french_agent",
|
||||
instructions="You respond in French. Always reply to the user's question in French.",
|
||||
)
|
||||
|
||||
# Create orchestrator with conditional tools
|
||||
orchestrator = Agent(
|
||||
name="orchestrator",
|
||||
instructions=(
|
||||
"You are a multilingual assistant. You use the tools given to you to respond to users. "
|
||||
"You must call ALL available tools to provide responses in different languages. "
|
||||
"You never respond in languages yourself, you always use the provided tools."
|
||||
),
|
||||
tools=[
|
||||
spanish_agent.as_tool(
|
||||
tool_name="respond_spanish",
|
||||
tool_description="Respond to the user's question in Spanish",
|
||||
is_enabled=True, # Always enabled
|
||||
),
|
||||
french_agent.as_tool(
|
||||
tool_name="respond_french",
|
||||
tool_description="Respond to the user's question in French",
|
||||
is_enabled=french_enabled,
|
||||
),
|
||||
],
|
||||
)
|
||||
|
||||
async def main():
|
||||
context = RunContextWrapper(LanguageContext(language_preference="french_spanish"))
|
||||
result = await Runner.run(orchestrator, "How are you?", context=context.context)
|
||||
print(result.final_output)
|
||||
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
`is_enabled` パラメーターは次を受け付けます。
|
||||
|
||||
- **ブール値**: `True`(常に有効)または `False`(常に無効)
|
||||
- **呼び出し可能関数**: `(context, agent)` を受け取り、ブール値を返す関数
|
||||
- **非同期関数**: 複雑な条件ロジック用の非同期関数
|
||||
|
||||
無効化されたツールは、ランタイムで LLM から完全に隠されるため、次のような用途に便利です。
|
||||
|
||||
- ユーザー権限に基づく機能ゲーティング
|
||||
- 環境固有のツール可用性(dev と prod)
|
||||
- 異なるツール構成の A/B テスト
|
||||
- ランタイム状態に基づく動的なツールフィルタリング
|
||||
|
||||
## 実験的機能: Codex ツール
|
||||
|
||||
`codex_tool` は Codex CLI をラップし、エージェントがツール呼び出し中にワークスペーススコープのタスク(シェル、ファイル編集、MCP ツール)を実行できるようにします。このサーフェスは実験的であり、変更される可能性があります。
|
||||
|
||||
メインエージェントが現在の実行を離れることなく、範囲を限定したワークスペースタスクを Codex に委任したい場合に使用します。デフォルトでは、ツール名は `codex` です。カスタム名を設定する場合、それは `codex` であるか、`codex_` で始まる必要があります。エージェントに複数の Codex ツールが含まれる場合、それぞれが一意の名前を使用する必要があります。
|
||||
|
||||
```python
|
||||
from agents import Agent
|
||||
from agents.extensions.experimental.codex import ThreadOptions, TurnOptions, codex_tool
|
||||
|
||||
agent = Agent(
|
||||
name="Codex Agent",
|
||||
instructions="Use the codex tool to inspect the workspace and answer the question.",
|
||||
tools=[
|
||||
codex_tool(
|
||||
sandbox_mode="workspace-write",
|
||||
working_directory="/path/to/repo",
|
||||
default_thread_options=ThreadOptions(
|
||||
model="gpt-5.5",
|
||||
model_reasoning_effort="low",
|
||||
network_access_enabled=True,
|
||||
web_search_mode="disabled",
|
||||
approval_policy="never",
|
||||
),
|
||||
default_turn_options=TurnOptions(
|
||||
idle_timeout_seconds=60,
|
||||
),
|
||||
persist_session=True,
|
||||
)
|
||||
],
|
||||
)
|
||||
```
|
||||
|
||||
次のオプショングループから始めてください。
|
||||
|
||||
- 実行サーフェス: `sandbox_mode` と `working_directory` は、Codex が操作できる場所を定義します。これらは組み合わせて使用し、作業ディレクトリが Git リポジトリ内にない場合は `skip_git_repo_check=True` を設定してください。
|
||||
- スレッドデフォルト: `default_thread_options=ThreadOptions(...)` は、モデル、推論エフォート、承認ポリシー、追加ディレクトリ、ネットワークアクセス、Web 検索モードを構成します。レガシーな `web_search_enabled` よりも `web_search_mode` を優先してください。
|
||||
- ターンデフォルト: `default_turn_options=TurnOptions(...)` は、`idle_timeout_seconds` や任意のキャンセル `signal` など、ターンごとの動作を構成します。
|
||||
- ツール I/O: ツール呼び出しには、`{ "type": "text", "text": ... }` または `{ "type": "local_image", "path": ... }` を持つ `inputs` 項目が少なくとも 1 つ含まれている必要があります。`output_schema` により、構造化された Codex 応答を要求できます。
|
||||
|
||||
スレッド再利用と永続化は別々の制御です。
|
||||
|
||||
- `persist_session=True` は、同じツールインスタンスへの繰り返し呼び出しに 1 つの Codex スレッドを再利用します。
|
||||
- `use_run_context_thread_id=True` は、同じ可変コンテキストオブジェクトを共有する実行間で、実行コンテキスト内にスレッド ID を保存して再利用します。
|
||||
- スレッド ID の優先順位は、呼び出しごとの `thread_id`、次に実行コンテキストのスレッド ID(有効な場合)、次に構成済みの `thread_id` オプションです。
|
||||
- デフォルトの実行コンテキストキーは、`name="codex"` の場合は `codex_thread_id`、`name="codex_<suffix>"` の場合は `codex_thread_id_<suffix>` です。`run_context_thread_id_key` で上書きできます。
|
||||
|
||||
ランタイム構成:
|
||||
|
||||
- 認証: `CODEX_API_KEY`(推奨)または `OPENAI_API_KEY` を設定するか、`codex_options={"api_key": "..."}` を渡します。
|
||||
- ランタイム: `codex_options.base_url` は CLI ベース URL を上書きします。
|
||||
- バイナリ解決: CLI パスを固定するには、`codex_options.codex_path_override`(または `CODEX_PATH`)を設定します。そうでない場合、SDK は `PATH` から `codex` を解決し、その後、同梱のベンダーバイナリにフォールバックします。
|
||||
- 環境: `codex_options.env` はサブプロセス環境を完全に制御します。これが指定された場合、サブプロセスは `os.environ` を継承しません。
|
||||
- ストリーム制限: `codex_options.codex_subprocess_stream_limit_bytes`(または `OPENAI_AGENTS_CODEX_SUBPROCESS_STREAM_LIMIT_BYTES`)は stdout/stderr リーダー制限を制御します。有効範囲は `65536` から `67108864` で、デフォルトは `8388608` です。
|
||||
- ストリーミング: `on_stream` はスレッド/ターンのライフサイクルイベントと項目イベント(`reasoning`、`command_execution`、`mcp_tool_call`、`file_change`、`web_search`、`todo_list`、`error` の項目更新)を受け取ります。
|
||||
- 出力: 実行結果には `response`、`usage`、`thread_id` が含まれます。usage は `RunContextWrapper.usage` に追加されます。
|
||||
|
||||
リファレンス:
|
||||
|
||||
- [Codex ツール API リファレンス](ref/extensions/experimental/codex/codex_tool.md)
|
||||
- [ThreadOptions リファレンス](ref/extensions/experimental/codex/thread_options.md)
|
||||
- [TurnOptions リファレンス](ref/extensions/experimental/codex/turn_options.md)
|
||||
- 完全に実行可能なサンプルについては、`examples/tools/codex.py` と `examples/tools/codex_same_thread.py` を参照してください。
|
||||
+161
-52
@@ -1,51 +1,103 @@
|
||||
---
|
||||
search:
|
||||
exclude: true
|
||||
---
|
||||
# トレーシング
|
||||
|
||||
Agents SDK にはトレーシング機能が組み込まれており、エージェントの実行中に発生するイベント( LLM 生成、ツール呼び出し、ハンドオフ、ガードレール、カスタムイベントなど)の包括的な記録を収集します。[Traces ダッシュボード](https://platform.openai.com/traces) を使用することで、開発中や本番環境でワークフローのデバッグ、可視化、監視が可能です。
|
||||
Agents SDK には組み込みのトレーシングが含まれており、エージェント実行中のイベントの包括的な記録を収集します。これには、LLM 生成、ツール呼び出し、ハンドオフ、ガードレール、さらには発生するカスタムイベントまで含まれます。[Traces ダッシュボード](https://platform.openai.com/traces)を使用すると、開発中および本番環境でワークフローをデバッグ、可視化、監視できます。
|
||||
|
||||
!!!note
|
||||
|
||||
トレーシングはデフォルトで有効になっています。トレーシングを無効にする方法は 2 つあります:
|
||||
トレーシングはデフォルトで有効になっています。一般的には、次の 3 つの方法で無効にできます。
|
||||
|
||||
1. 環境変数 `OPENAI_AGENTS_DISABLE_TRACING=1` を設定することで、グローバルにトレーシングを無効化できます。
|
||||
2. 単一の実行に対しては [`agents.run.RunConfig.tracing_disabled`][] を `True` に設定することで無効化できます。
|
||||
1. 環境変数 `OPENAI_AGENTS_DISABLE_TRACING=1` を設定することで、トレーシングをグローバルに無効化できます
|
||||
2. [`set_tracing_disabled(True)`][agents.set_tracing_disabled] を使って、コード内でトレーシングをグローバルに無効化できます
|
||||
3. [`agents.run.RunConfig.tracing_disabled`][] を `True` に設定することで、単一の実行についてトレーシングを無効化できます
|
||||
|
||||
***OpenAI の API を使用し、Zero Data Retention (ZDR) ポリシーの下で運用している組織では、トレーシングは利用できません。***
|
||||
***OpenAI の API を使用し、Zero Data Retention (ZDR) ポリシー下で運用している組織では、トレーシングは利用できません。***
|
||||
|
||||
## トレースとスパン
|
||||
|
||||
- **トレース** は 1 つの「ワークフロー」のエンドツーエンドの操作を表します。トレースはスパンで構成されます。トレースには以下のプロパティがあります:
|
||||
- `workflow_name`: 論理的なワークフローやアプリ名です。例: "Code generation" や "Customer service" など。
|
||||
- `trace_id`: トレースの一意な ID です。指定しない場合は自動生成されます。フォーマットは `trace_<32_alphanumeric>` である必要があります。
|
||||
- `group_id`: オプションのグループ ID で、同じ会話からの複数のトレースをリンクするために使用します。例: チャットスレッド ID など。
|
||||
- `disabled`: True の場合、このトレースは記録されません。
|
||||
- `metadata`: トレースに付加するオプションのメタデータです。
|
||||
- **スパン** は開始時刻と終了時刻を持つ操作を表します。スパンには以下があります:
|
||||
- `started_at` および `ended_at` タイムスタンプ
|
||||
- 所属するトレースを示す `trace_id`
|
||||
- このスパンの親スパンを指す `parent_id`(存在する場合)
|
||||
- スパンに関する情報を含む `span_data`。例: `AgentSpanData` はエージェントに関する情報、`GenerationSpanData` は LLM 生成に関する情報など。
|
||||
- **トレース** は、「ワークフロー」の単一のエンドツーエンド操作を表します。これらはスパンで構成されます。トレースには以下のプロパティがあります:
|
||||
- `workflow_name`: 論理的なワークフローまたはアプリです。たとえば「コード生成」や「カスタマーサービス」です。
|
||||
- `trace_id`: トレースの一意の ID です。渡さない場合は自動生成されます。形式は `trace_<32_alphanumeric>` でなければなりません。
|
||||
- `group_id`: 任意のグループ ID で、同じ会話からの複数のトレースを関連付けるために使用します。たとえば、チャットスレッド ID を使用できます。
|
||||
- `disabled`: True の場合、トレースは記録されません。
|
||||
- `metadata`: トレースの任意のメタデータです。
|
||||
- **スパン** は、開始時刻と終了時刻を持つ操作を表します。スパンには以下があります:
|
||||
- `started_at` と `ended_at` のタイムスタンプ。
|
||||
- `trace_id`: そのスパンが属するトレースを表します
|
||||
- `parent_id`: このスパンの親スパン (存在する場合) を指します
|
||||
- `span_data`: スパンに関する情報です。たとえば、`AgentSpanData` にはエージェントに関する情報が含まれ、`GenerationSpanData` には LLM 生成に関する情報が含まれる、などです。
|
||||
|
||||
## デフォルトのトレーシング
|
||||
|
||||
デフォルトでは、 SDK は以下をトレースします:
|
||||
デフォルトでは、SDK は以下をトレースします:
|
||||
|
||||
- `Runner.{run, run_sync, run_streamed}()` 全体が `trace()` でラップされます。
|
||||
- エージェントが実行されるたびに `agent_span()` でラップされます。
|
||||
- LLM 生成は `generation_span()` でラップされます。
|
||||
- 関数ツール呼び出しはそれぞれ `function_span()` でラップされます。
|
||||
- ガードレールは `guardrail_span()` でラップされます。
|
||||
- ハンドオフは `handoff_span()` でラップされます。
|
||||
- 音声入力(音声からテキスト)は `transcription_span()` でラップされます。
|
||||
- 音声出力(テキストから音声)は `speech_span()` でラップされます。
|
||||
- 関連する音声スパンは `speech_group_span()` の下にまとめられる場合があります。
|
||||
- エージェントが実行されるたびに、`agent_span()` でラップされます
|
||||
- LLM 生成は `generation_span()` でラップされます
|
||||
- 関数ツール呼び出しはそれぞれ `function_span()` でラップされます
|
||||
- ガードレールは `guardrail_span()` でラップされます
|
||||
- ハンドオフは `handoff_span()` でラップされます
|
||||
- 音声入力 (音声テキスト変換) は `transcription_span()` でラップされます
|
||||
- 音声出力 (テキスト音声変換) は `speech_span()` でラップされます
|
||||
- 関連する音声スパンは `speech_group_span()` の下に親子関係として配置される場合があります
|
||||
|
||||
デフォルトでは、トレース名は "Agent trace" です。`trace` を使用する場合、この名前を設定できます。また、[`RunConfig`][agents.run.RunConfig] で名前やその他のプロパティを設定することも可能です。
|
||||
デフォルトでは、トレース名は "Agent workflow" です。`trace` を使用する場合はこの名前を設定できます。また、[`RunConfig`][agents.run.RunConfig] で名前やその他のプロパティを構成できます。
|
||||
|
||||
さらに、[カスタムトレースプロセッサー](#custom-tracing-processors) を設定して、トレースを他の宛先に送信することもできます(置き換えや追加の宛先として)。
|
||||
さらに、[カスタムトレーシングプロセッサー](#custom-tracing-processors)を設定して、トレースを他の送信先へ送信することもできます (置き換え、または副次的な送信先として)。
|
||||
|
||||
## より高レベルのトレース
|
||||
## 長時間実行ワーカーと即時エクスポート
|
||||
|
||||
複数回の `run()` 呼び出しを 1 つのトレースにまとめたい場合があります。その場合、コード全体を `trace()` でラップしてください。
|
||||
デフォルトの [`BatchTraceProcessor`][agents.tracing.processors.BatchTraceProcessor] は、数秒ごとにバックグラウンドでトレースをエクスポートします。
|
||||
または、メモリ内キューがサイズトリガーに達した場合はそれより早くエクスポートし、
|
||||
プロセス終了時にも最終的なフラッシュを実行します。Celery、RQ、Dramatiq、FastAPI バックグラウンドタスクのような長時間実行ワーカーでは、通常、追加コードなしでトレースが自動的にエクスポートされますが、各ジョブの完了直後に Traces ダッシュボードへ表示されない場合があります。
|
||||
|
||||
作業単位の終了時に即時配信の保証が必要な場合は、トレースコンテキストが終了した後に
|
||||
[`flush_traces()`][agents.tracing.flush_traces] を呼び出してください。
|
||||
|
||||
```python
|
||||
from agents import Runner, flush_traces, trace
|
||||
|
||||
|
||||
@celery_app.task
|
||||
def run_agent_task(prompt: str):
|
||||
try:
|
||||
with trace("celery_task"):
|
||||
result = Runner.run_sync(agent, prompt)
|
||||
return result.final_output
|
||||
finally:
|
||||
flush_traces()
|
||||
```
|
||||
|
||||
```python
|
||||
from fastapi import BackgroundTasks, FastAPI
|
||||
from agents import Runner, flush_traces, trace
|
||||
|
||||
app = FastAPI()
|
||||
|
||||
|
||||
def process_in_background(prompt: str) -> None:
|
||||
try:
|
||||
with trace("background_job"):
|
||||
Runner.run_sync(agent, prompt)
|
||||
finally:
|
||||
flush_traces()
|
||||
|
||||
|
||||
@app.post("/run")
|
||||
async def run(prompt: str, background_tasks: BackgroundTasks):
|
||||
background_tasks.add_task(process_in_background, prompt)
|
||||
return {"status": "queued"}
|
||||
```
|
||||
|
||||
[`flush_traces()`][agents.tracing.flush_traces] は、現在バッファリングされているトレースとスパンが
|
||||
エクスポートされるまでブロックします。そのため、部分的に構築されたトレースをフラッシュしないように、`trace()` が閉じた後に呼び出してください。デフォルトのエクスポート遅延で許容できる場合は、この呼び出しを省略できます。
|
||||
|
||||
## 上位レベルのトレース
|
||||
|
||||
場合によっては、`run()` への複数回の呼び出しを 1 つのトレースの一部にしたいことがあります。コード全体を `trace()` でラップすることで実現できます。
|
||||
|
||||
```python
|
||||
from agents import Agent, Runner, trace
|
||||
@@ -60,57 +112,114 @@ async def main():
|
||||
print(f"Rating: {second_result.final_output}")
|
||||
```
|
||||
|
||||
1. 2 回の `Runner.run` 呼び出しが `with trace()` でラップされているため、個々の実行は 2 つのトレースを作成するのではなく、全体のトレースの一部となります。
|
||||
1. 2 つの `Runner.run` 呼び出しは `with trace()` でラップされているため、個別の実行は 2 つのトレースを作成するのではなく、全体のトレースの一部になります。
|
||||
|
||||
## トレースの作成
|
||||
|
||||
[`trace()`][agents.tracing.trace] 関数を使ってトレースを作成できます。トレースは開始と終了が必要です。方法は 2 つあります:
|
||||
[`trace()`][agents.tracing.trace] 関数を使用してトレースを作成できます。トレースは開始および終了する必要があります。これには 2 つの方法があります:
|
||||
|
||||
1. **推奨**: トレースをコンテキストマネージャーとして使用します。例: `with trace(...) as my_trace`。これにより、トレースの開始と終了が自動的に行われます。
|
||||
2. [`trace.start()`][agents.tracing.Trace.start] および [`trace.finish()`][agents.tracing.Trace.finish] を手動で呼び出すこともできます。
|
||||
1. **推奨**: トレースをコンテキストマネージャーとして使用します。つまり、`with trace(...) as my_trace` とします。これにより、適切なタイミングでトレースが自動的に開始および終了されます。
|
||||
2. [`trace.start()`][agents.tracing.Trace.start] と [`trace.finish()`][agents.tracing.Trace.finish] を手動で呼び出すこともできます。
|
||||
|
||||
現在のトレースは Python の [`contextvar`](https://docs.python.org/3/library/contextvars.html) で管理されています。これにより、自動的に並行処理にも対応します。トレースを手動で開始・終了する場合は、`start()`/`finish()` に `mark_as_current` および `reset_current` を渡して現在のトレースを更新する必要があります。
|
||||
現在のトレースは Python の [`contextvar`](https://docs.python.org/3/library/contextvars.html) によって追跡されます。これにより、並行処理でも自動的に機能します。トレースを手動で開始/終了する場合は、現在のトレースを更新するために `start()`/`finish()` に `mark_as_current` と `reset_current` を渡す必要があります。
|
||||
|
||||
## スパンの作成
|
||||
|
||||
さまざまな [`*_span()`][agents.tracing.create] メソッドを使ってスパンを作成できます。通常、スパンを手動で作成する必要はありません。カスタムスパン情報を追跡するための [`custom_span()`][agents.tracing.custom_span] 関数も用意されています。
|
||||
各種 [`*_span()`][agents.tracing.create] メソッドを使用してスパンを作成できます。一般に、スパンを手動で作成する必要はありません。カスタムスパン情報を追跡するために [`custom_span()`][agents.tracing.custom_span] 関数が利用できます。
|
||||
|
||||
スパンは自動的に現在のトレースの一部となり、最も近い現在のスパンの下にネストされます。これは Python の [`contextvar`](https://docs.python.org/3/library/contextvars.html) で管理されています。
|
||||
スパンは自動的に現在のトレースの一部となり、現在の最も近いスパンの下にネストされます。この現在のスパンは Python の [`contextvar`](https://docs.python.org/3/library/contextvars.html) によって追跡されます。
|
||||
|
||||
## 機微なデータ
|
||||
## 機密データ
|
||||
|
||||
一部のスパンは、機微なデータを記録する場合があります。
|
||||
一部のスパンは、機密性の高い可能性があるデータをキャプチャする場合があります。
|
||||
|
||||
`generation_span()` は LLM 生成の入力・出力を保存し、`function_span()` は関数呼び出しの入力・出力を保存します。これらには機微なデータが含まれる場合があるため、[`RunConfig.trace_include_sensitive_data`][agents.run.RunConfig.trace_include_sensitive_data] でそのデータの記録を無効化できます。
|
||||
`generation_span()` は LLM 生成の入力/出力を保存し、`function_span()` は関数呼び出しの入力/出力を保存します。これらには機密データが含まれる可能性があるため、[`RunConfig.trace_include_sensitive_data`][agents.run.RunConfig.trace_include_sensitive_data] を使ってそのデータのキャプチャを無効化できます。
|
||||
|
||||
同様に、音声スパンはデフォルトで入力・出力音声の base64 エンコード PCM データを含みます。[`VoicePipelineConfig.trace_include_sensitive_audio_data`][agents.voice.pipeline_config.VoicePipelineConfig.trace_include_sensitive_audio_data] を設定することで、この音声データの記録を無効化できます。
|
||||
同様に、音声スパンには、デフォルトで入力および出力音声の base64 エンコードされた PCM データが含まれます。この音声データのキャプチャは、[`VoicePipelineConfig.trace_include_sensitive_audio_data`][agents.voice.pipeline_config.VoicePipelineConfig.trace_include_sensitive_audio_data] を構成することで無効化できます。
|
||||
|
||||
## カスタムトレースプロセッサー
|
||||
デフォルトでは、`trace_include_sensitive_data` は `True` です。アプリを実行する前に `OPENAI_AGENTS_TRACE_INCLUDE_SENSITIVE_DATA` 環境変数を `true/1` または `false/0` にエクスポートすることで、コードなしでデフォルトを設定できます。
|
||||
|
||||
トレーシングの高レベルなアーキテクチャは以下の通りです:
|
||||
## カスタムトレーシングプロセッサー
|
||||
|
||||
- 初期化時にグローバルな [`TraceProvider`][agents.tracing.setup.TraceProvider] を作成し、トレースの生成を担当します。
|
||||
- `TraceProvider` は [`BatchTraceProcessor`][agents.tracing.processors.BatchTraceProcessor] で構成されており、トレースやスパンをバッチで [`BackendSpanExporter`][agents.tracing.processors.BackendSpanExporter] に送信します。これにより、スパンやトレースが OpenAI バックエンドにバッチでエクスポートされます。
|
||||
トレーシングの高レベルアーキテクチャは次のとおりです:
|
||||
|
||||
このデフォルト設定をカスタマイズし、トレースを別のバックエンドや追加のバックエンドに送信したり、エクスポーターの動作を変更したりするには、2 つの方法があります:
|
||||
- 初期化時に、トレースの作成を担当するグローバルな [`TraceProvider`][agents.tracing.setup.TraceProvider] を作成します。
|
||||
- `TraceProvider` は、トレース/スパンをバッチで [`BackendSpanExporter`][agents.tracing.processors.BackendSpanExporter] に送信する [`BatchTraceProcessor`][agents.tracing.processors.BatchTraceProcessor] で構成します。`BackendSpanExporter` は、スパンとトレースをバッチで OpenAI バックエンドにエクスポートします。
|
||||
|
||||
1. [`add_trace_processor()`][agents.tracing.add_trace_processor] を使うと、**追加の** トレースプロセッサーを追加できます。これにより、トレースやスパンが準備できた時点で独自の処理を行うことができ、OpenAI バックエンドへの送信に加えて独自の処理が可能です。
|
||||
2. [`set_trace_processors()`][agents.tracing.set_trace_processors] を使うと、デフォルトのプロセッサーを独自のトレースプロセッサーに**置き換える**ことができます。この場合、OpenAI バックエンドにトレースが送信されるのは、`TracingProcessor` を含めた場合のみです。
|
||||
このデフォルト設定をカスタマイズして、代替または追加のバックエンドへトレースを送信したり、エクスポーターの動作を変更したりするには、2 つの選択肢があります:
|
||||
|
||||
## 外部トレースプロセッサー一覧
|
||||
1. [`add_trace_processor()`][agents.tracing.add_trace_processor] を使用すると、利用可能になったトレースとスパンを受け取る **追加** のトレーシングプロセッサーを追加できます。これにより、トレースを OpenAI バックエンドに送信することに加えて、独自の処理を実行できます。
|
||||
2. [`set_trace_processors()`][agents.tracing.set_trace_processors] を使用すると、デフォルトのプロセッサーを独自のトレーシングプロセッサーで **置き換える** ことができます。これは、それを行う `TracingProcessor` を含めない限り、トレースが OpenAI バックエンドに送信されないことを意味します。
|
||||
|
||||
|
||||
## 非 OpenAI モデルでのトレーシング
|
||||
|
||||
非 OpenAI モデルで OpenAI API キーを使用すると、トレーシングを無効にする必要なく、OpenAI Traces ダッシュボードで無料のトレーシングを有効にできます。アダプターの選択とセットアップ時の注意点については、モデルガイドの[サードパーティアダプター](models/index.md#third-party-adapters)セクションを参照してください。
|
||||
|
||||
```python
|
||||
import os
|
||||
from agents import set_tracing_export_api_key, Agent, Runner
|
||||
from agents.extensions.models.any_llm_model import AnyLLMModel
|
||||
|
||||
tracing_api_key = os.environ["OPENAI_API_KEY"]
|
||||
set_tracing_export_api_key(tracing_api_key)
|
||||
|
||||
model = AnyLLMModel(
|
||||
model="your-provider/your-model-name",
|
||||
api_key="your-api-key",
|
||||
)
|
||||
|
||||
agent = Agent(
|
||||
name="Assistant",
|
||||
model=model,
|
||||
)
|
||||
```
|
||||
|
||||
単一の実行で異なるトレーシングキーだけが必要な場合は、グローバルエクスポーターを変更する代わりに `RunConfig` 経由で渡してください。
|
||||
|
||||
```python
|
||||
from agents import Runner, RunConfig
|
||||
|
||||
await Runner.run(
|
||||
agent,
|
||||
input="Hello",
|
||||
run_config=RunConfig(tracing={"api_key": "sk-tracing-123"}),
|
||||
)
|
||||
```
|
||||
|
||||
## 補足事項
|
||||
- Openai Traces ダッシュボードで無料のトレースを表示できます。
|
||||
|
||||
|
||||
## エコシステム統合
|
||||
|
||||
以下のコミュニティおよびベンダー統合は、OpenAI Agents SDK のトレーシングインターフェイスに対応しています。
|
||||
|
||||
### 外部トレーシングプロセッサー一覧
|
||||
|
||||
- [Weights & Biases](https://weave-docs.wandb.ai/guides/integrations/openai_agents)
|
||||
- [Arize-Phoenix](https://docs.arize.com/phoenix/tracing/integrations-tracing/openai-agents-sdk)
|
||||
- [MLflow (self-hosted/OSS](https://mlflow.org/docs/latest/tracing/integrations/openai-agent)
|
||||
- [MLflow (Databricks hosted](https://docs.databricks.com/aws/en/mlflow/mlflow-tracing#-automatic-tracing)
|
||||
- [Future AGI](https://docs.futureagi.com/future-agi/products/observability/auto-instrumentation/openai_agents)
|
||||
- [MLflow (セルフホスト/OSS)](https://mlflow.org/docs/latest/tracing/integrations/openai-agent)
|
||||
- [MLflow (Databricks ホスト)](https://docs.databricks.com/aws/en/mlflow/mlflow-tracing#-automatic-tracing)
|
||||
- [Braintrust](https://braintrust.dev/docs/guides/traces/integrations#openai-agents-sdk)
|
||||
- [Pydantic Logfire](https://logfire.pydantic.dev/docs/integrations/llms/openai/#openai-agents)
|
||||
- [AgentOps](https://docs.agentops.ai/v1/integrations/agentssdk)
|
||||
- [Scorecard](https://docs.scorecard.io/docs/documentation/features/tracing#openai-agents-sdk-integration)
|
||||
- [Keywords AI](https://docs.keywordsai.co/integration/development-frameworks/openai-agent)
|
||||
- [Respan](https://respan.ai/docs/integrations/tracing/openai-agents-sdk)
|
||||
- [LangSmith](https://docs.smith.langchain.com/observability/how_to_guides/trace_with_openai_agents_sdk)
|
||||
- [Maxim AI](https://www.getmaxim.ai/docs/observe/integrations/openai-agents-sdk)
|
||||
- [Comet Opik](https://www.comet.com/docs/opik/tracing/integrations/openai_agents)
|
||||
- [Langfuse](https://langfuse.com/docs/integrations/openaiagentssdk/openai-agents)
|
||||
- [Langtrace](https://docs.langtrace.ai/supported-integrations/llm-frameworks/openai-agents-sdk)
|
||||
- [Okahu-Monocle](https://github.com/monocle2ai/monocle)
|
||||
- [Okahu-Monocle](https://github.com/monocle2ai/monocle)
|
||||
- [Galileo](https://v2docs.galileo.ai/integrations/openai-agent-integration#openai-agent-integration)
|
||||
- [Portkey AI](https://portkey.ai/docs/integrations/agents/openai-agents)
|
||||
- [LangDB AI](https://docs.langdb.ai/getting-started/working-with-agent-frameworks/working-with-openai-agents-sdk)
|
||||
- [Agenta](https://docs.agenta.ai/observability/integrations/openai-agents)
|
||||
- [PostHog](https://posthog.com/docs/llm-analytics/installation/openai-agents)
|
||||
- [Traccia](https://traccia.ai/docs/integrations/openai-agents)
|
||||
- [PromptLayer](https://docs.promptlayer.com/languages/integrations#openai-agents-sdk)
|
||||
- [HoneyHive](https://docs.honeyhive.ai/v2/integrations/openai-agents)
|
||||
- [Asqav](https://www.asqav.com/docs/integrations#openai-agents)
|
||||
- [Datadog](https://docs.datadoghq.com/llm_observability/instrumentation/auto_instrumentation/?tab=python#openai-agents)
|
||||
@@ -0,0 +1,90 @@
|
||||
---
|
||||
search:
|
||||
exclude: true
|
||||
---
|
||||
# 使用量
|
||||
|
||||
Agents SDK は、各実行のトークン使用量を自動的に追跡します。実行コンテキストからアクセスでき、コストの監視、上限の適用、分析情報の記録に利用できます。
|
||||
|
||||
## 追跡対象
|
||||
|
||||
- **requests**: 実行された LLM API 呼び出しの数
|
||||
- **input_tokens**: 送信された入力トークンの合計
|
||||
- **output_tokens**: 受信された出力トークンの合計
|
||||
- **total_tokens**: 入力 + 出力
|
||||
- **request_usage_entries**: リクエストごとの使用量内訳のリスト
|
||||
- **details**:
|
||||
- `input_tokens_details.cached_tokens`
|
||||
- `output_tokens_details.reasoning_tokens`
|
||||
|
||||
## 実行からの使用量へのアクセス
|
||||
|
||||
`Runner.run(...)` の後に、`result.context_wrapper.usage` 経由で使用量にアクセスします。
|
||||
|
||||
```python
|
||||
result = await Runner.run(agent, "What's the weather in Tokyo?")
|
||||
usage = result.context_wrapper.usage
|
||||
|
||||
print("Requests:", usage.requests)
|
||||
print("Input tokens:", usage.input_tokens)
|
||||
print("Output tokens:", usage.output_tokens)
|
||||
print("Total tokens:", usage.total_tokens)
|
||||
```
|
||||
|
||||
使用量は、実行中のすべてのモデル呼び出し(ツール呼び出しやハンドオフを含む)にわたって集計されます。
|
||||
|
||||
### サードパーティアダプターでの使用量の有効化
|
||||
|
||||
使用量レポートは、サードパーティアダプターやプロバイダーのバックエンドによって異なります。アダプター経由のモデルに依存しており、正確な `result.context_wrapper.usage` 値が必要な場合は、次の点に注意してください。
|
||||
|
||||
- `AnyLLMModel` では、上流プロバイダーが使用量を返す場合、使用量は自動的に伝播されます。ストリーミングされた Chat Completions バックエンドでは、使用量チャンクが出力される前に `ModelSettings(include_usage=True)` が必要になる場合があります。
|
||||
- `LitellmModel` では、一部のプロバイダーのバックエンドがデフォルトで使用量をレポートしないため、`ModelSettings(include_usage=True)` が必要になることがよくあります。
|
||||
|
||||
Models ガイドの [サードパーティアダプター](models/index.md#third-party-adapters) セクションにあるアダプター固有の注記を確認し、デプロイ予定の正確なプロバイダーバックエンドを検証してください。
|
||||
|
||||
## リクエストごとの使用量追跡
|
||||
|
||||
SDK は、各 API リクエストの使用量を `request_usage_entries` で自動的に追跡します。これは詳細なコスト計算やコンテキストウィンドウ消費量の監視に役立ちます。
|
||||
|
||||
```python
|
||||
result = await Runner.run(agent, "What's the weather in Tokyo?")
|
||||
|
||||
for i, request in enumerate(result.context_wrapper.usage.request_usage_entries):
|
||||
print(f"Request {i + 1}: {request.input_tokens} in, {request.output_tokens} out")
|
||||
```
|
||||
|
||||
## セッションでの使用量へのアクセス
|
||||
|
||||
`Session`(例: `SQLiteSession`)を使用する場合、`Runner.run(...)` の各呼び出しは、その特定の実行の使用量を返します。セッションはコンテキスト用に会話履歴を保持しますが、各実行の使用量は独立しています。
|
||||
|
||||
```python
|
||||
session = SQLiteSession("my_conversation")
|
||||
|
||||
first = await Runner.run(agent, "Hi!", session=session)
|
||||
print(first.context_wrapper.usage.total_tokens) # Usage for first run
|
||||
|
||||
second = await Runner.run(agent, "Can you elaborate?", session=session)
|
||||
print(second.context_wrapper.usage.total_tokens) # Usage for second run
|
||||
```
|
||||
|
||||
セッションは実行間で会話コンテキストを保持しますが、各 `Runner.run()` 呼び出しで返される使用量メトリクスは、その特定の実行のみを表すことに注意してください。セッションでは、以前のメッセージが各実行の入力として再投入される場合があり、その結果、以降のターンで入力トークン数に影響します。
|
||||
|
||||
## フックでの使用量の利用
|
||||
|
||||
`RunHooks` を使用している場合、各フックに渡される `context` オブジェクトには `usage` が含まれます。これにより、主要なライフサイクルのタイミングで使用量をログに記録できます。
|
||||
|
||||
```python
|
||||
class MyHooks(RunHooks):
|
||||
async def on_agent_end(self, context: RunContextWrapper, agent: Agent, output: Any) -> None:
|
||||
u = context.usage
|
||||
print(f"{agent.name} → {u.requests} requests, {u.total_tokens} total tokens")
|
||||
```
|
||||
|
||||
## API リファレンス
|
||||
|
||||
詳細な API ドキュメントについては、以下を参照してください。
|
||||
|
||||
- [`Usage`][agents.usage.Usage] - 使用量追跡データ構造
|
||||
- [`RequestUsage`][agents.usage.RequestUsage] - リクエストごとの使用量詳細
|
||||
- [`RunContextWrapper`][agents.run.RunContextWrapper] - 実行コンテキストから使用量にアクセス
|
||||
- [`RunHooks`][agents.run.RunHooks] - 使用量追跡ライフサイクルへのフックイン
|
||||
+41
-16
@@ -1,10 +1,14 @@
|
||||
---
|
||||
search:
|
||||
exclude: true
|
||||
---
|
||||
# エージェントの可視化
|
||||
|
||||
エージェントの可視化では、 **Graphviz** を使用してエージェントとその関係性を構造的にグラフィカルに表現できます。これは、アプリケーション内でエージェント、ツール、ハンドオフがどのように相互作用しているかを理解するのに役立ちます。
|
||||
エージェントの可視化では、**Graphviz** を使用して、エージェントとその関係を構造化されたグラフィカルな表現として生成できます。これは、アプリケーション内でエージェント、ツール、ハンドオフがどのように相互作用するかを理解するのに役立ちます。
|
||||
|
||||
## インストール
|
||||
|
||||
オプションの `viz` 依存グループをインストールします:
|
||||
任意の `viz` 依存関係グループをインストールします。
|
||||
|
||||
```bash
|
||||
pip install "openai-agents[viz]"
|
||||
@@ -12,16 +16,20 @@ pip install "openai-agents[viz]"
|
||||
|
||||
## グラフの生成
|
||||
|
||||
`draw_graph` 関数を使ってエージェントの可視化を生成できます。この関数は、以下のような有向グラフを作成します:
|
||||
`draw_graph` 関数を使用して、エージェントの可視化を生成できます。この関数は、以下のような有向グラフを作成します。
|
||||
|
||||
- **エージェント** は黄色のボックスで表されます。
|
||||
- **MCP サーバー** は灰色のボックスで表されます。
|
||||
- **ツール** は緑色の楕円で表されます。
|
||||
- **ハンドオフ** は、あるエージェントから別のエージェントへの有向エッジで表されます。
|
||||
- **ハンドオフ** は、あるエージェントから別のエージェントへの有向エッジとして表されます。
|
||||
|
||||
### 使用例
|
||||
|
||||
```python
|
||||
import os
|
||||
|
||||
from agents import Agent, function_tool
|
||||
from agents.mcp.server import MCPServerStdio
|
||||
from agents.extensions.visualization import draw_graph
|
||||
|
||||
@function_tool
|
||||
@@ -38,47 +46,64 @@ english_agent = Agent(
|
||||
instructions="You only speak English",
|
||||
)
|
||||
|
||||
current_dir = os.path.dirname(os.path.abspath(__file__))
|
||||
samples_dir = os.path.join(current_dir, "sample_files")
|
||||
mcp_server = MCPServerStdio(
|
||||
name="Filesystem Server, via npx",
|
||||
params={
|
||||
"command": "npx",
|
||||
"args": ["-y", "@modelcontextprotocol/server-filesystem", samples_dir],
|
||||
},
|
||||
)
|
||||
|
||||
triage_agent = Agent(
|
||||
name="Triage agent",
|
||||
instructions="Handoff to the appropriate agent based on the language of the request.",
|
||||
handoffs=[spanish_agent, english_agent],
|
||||
tools=[get_weather],
|
||||
mcp_servers=[mcp_server],
|
||||
)
|
||||
|
||||
draw_graph(triage_agent)
|
||||
```
|
||||
|
||||

|
||||

|
||||
|
||||
これにより、 **triage agent** の構造とサブエージェントやツールとの接続が視覚的に表現されたグラフが生成されます。
|
||||
これにより、**トリアージエージェント** の構造と、サブエージェントおよびツールとの接続を視覚的に表すグラフが生成されます。
|
||||
|
||||
|
||||
## 可視化の理解
|
||||
|
||||
生成されたグラフには以下が含まれます:
|
||||
生成されるグラフには、以下が含まれます。
|
||||
|
||||
- エントリーポイントを示す **start node** (`__start__`)
|
||||
- 黄色で塗りつぶされた **長方形** で表されるエージェント
|
||||
- 緑色で塗りつぶされた **楕円** で表されるツール
|
||||
- エントリーポイントを示す **開始ノード** (`__start__`)。
|
||||
- 黄色で塗りつぶされた **長方形** として表されるエージェント。
|
||||
- 緑色で塗りつぶされた **楕円** として表されるツール。
|
||||
- 灰色で塗りつぶされた **長方形** として表される MCP サーバー。
|
||||
- 相互作用を示す有向エッジ:
|
||||
- エージェント間のハンドオフには **実線の矢印**
|
||||
- ツール呼び出しには **点線の矢印**
|
||||
- 実行が終了する場所を示す **end node** (`__end__`)
|
||||
- エージェント間のハンドオフを表す **実線矢印**。
|
||||
- ツール呼び出しを表す **点線矢印**。
|
||||
- MCP サーバー呼び出しを表す **破線矢印**。
|
||||
- 実行が終了する場所を示す **終了ノード** (`__end__`)。
|
||||
|
||||
**注:** MCP サーバーは、最近のバージョンの
|
||||
`agents` パッケージで描画されます(**v0.2.8** で確認済み)。可視化で MCP ボックスが表示されない場合は、
|
||||
最新リリースにアップグレードしてください。
|
||||
|
||||
## グラフのカスタマイズ
|
||||
|
||||
### グラフの表示
|
||||
デフォルトでは、`draw_graph` はグラフをインラインで表示します。グラフを別ウィンドウで表示したい場合は、次のように記述します:
|
||||
デフォルトでは、`draw_graph` はグラフをインラインで表示します。グラフを別ウィンドウで表示するには、次のように記述します。
|
||||
|
||||
```python
|
||||
draw_graph(triage_agent).view()
|
||||
```
|
||||
|
||||
### グラフの保存
|
||||
デフォルトでは、`draw_graph` はグラフをインラインで表示します。ファイルとして保存したい場合は、ファイル名を指定します:
|
||||
デフォルトでは、`draw_graph` はグラフをインラインで表示します。ファイルとして保存するには、ファイル名を指定します。
|
||||
|
||||
```python
|
||||
draw_graph(triage_agent, filename="agent_graph.png")
|
||||
draw_graph(triage_agent, filename="agent_graph")
|
||||
```
|
||||
|
||||
これにより、作業ディレクトリに `agent_graph.png` が生成されます。
|
||||
+22
-18
@@ -1,6 +1,10 @@
|
||||
---
|
||||
search:
|
||||
exclude: true
|
||||
---
|
||||
# パイプラインとワークフロー
|
||||
|
||||
[`VoicePipeline`][agents.voice.pipeline.VoicePipeline] は、エージェントワークフローを音声アプリに簡単に変換できるクラスです。実行したいワークフローを渡すと、パイプラインが入力音声の文字起こし、音声終了の検出、適切なタイミングでのワークフロー呼び出し、ワークフロー出力の音声化までを自動で処理します。
|
||||
[`VoicePipeline`][agents.voice.pipeline.VoicePipeline] は、エージェントを活用したワークフローを音声アプリに変換しやすくするクラスです。実行するワークフローを渡すと、パイプラインが入力音声の文字起こし、音声の終了検出、適切なタイミングでのワークフロー呼び出し、ワークフローの出力を音声に戻す処理を担います。
|
||||
|
||||
```mermaid
|
||||
graph LR
|
||||
@@ -30,29 +34,29 @@ graph LR
|
||||
|
||||
## パイプラインの設定
|
||||
|
||||
パイプラインを作成する際、以下の項目を設定できます。
|
||||
パイプラインを作成する際に、いくつかの項目を設定できます。
|
||||
|
||||
1. [`workflow`][agents.voice.workflow.VoiceWorkflowBase]:新しい音声が文字起こしされるたびに実行されるコードです。
|
||||
2. 使用する [`speech-to-text`][agents.voice.model.STTModel] および [`text-to-speech`][agents.voice.model.TTSModel] モデル
|
||||
3. [`config`][agents.voice.pipeline_config.VoicePipelineConfig]:以下のような設定が可能です。
|
||||
- モデルプロバイダー:モデル名をモデルにマッピングできます
|
||||
- トレーシング:トレーシングの有効/無効、音声ファイルのアップロード有無、ワークフロー名、トレース ID など
|
||||
- TTS および STT モデルの設定:プロンプト、言語、使用するデータ型など
|
||||
1. [`workflow`][agents.voice.workflow.VoiceWorkflowBase]。新しい音声が文字起こしされるたびに実行されるコードです。
|
||||
2. 使用する [`speech-to-text`][agents.voice.model.STTModel] モデルと [`text-to-speech`][agents.voice.model.TTSModel] モデル
|
||||
3. [`config`][agents.voice.pipeline_config.VoicePipelineConfig]。次のような項目を設定できます。
|
||||
- モデル名をモデルに対応付けられるモデルプロバイダー
|
||||
- トレーシング。トレーシングを無効にするか、音声ファイルをアップロードするか、ワークフロー名、トレース ID などを含みます。
|
||||
- TTS モデルと STT モデルの設定。プロンプト、言語、使用するデータ型などです。
|
||||
|
||||
## パイプラインの実行
|
||||
|
||||
パイプラインは [`run()`][agents.voice.pipeline.VoicePipeline.run] メソッドで実行できます。音声入力は 2 つの形式で渡せます。
|
||||
[`run()`][agents.voice.pipeline.VoicePipeline.run] メソッドを使ってパイプラインを実行できます。このメソッドでは、次の 2 つの形式で音声入力を渡せます。
|
||||
|
||||
1. [`AudioInput`][agents.voice.input.AudioInput]:完全な音声トランスクリプトがある場合に使用し、その内容に対する結果のみを生成します。話者が話し終えたタイミングを検出する必要がない場合(例:事前録音音声や push-to-talk アプリなど、ユーザーが話し終えたことが明確な場合)に便利です。
|
||||
2. [`StreamedAudioInput`][agents.voice.input.StreamedAudioInput]:ユーザーが話し終えたタイミングを検出する必要がある場合に使用します。音声チャンクを検出ごとにプッシュでき、VoicePipeline が「アクティビティ検出」と呼ばれるプロセスを通じて、適切なタイミングでエージェントワークフローを自動実行します。
|
||||
1. [`AudioInput`][agents.voice.input.AudioInput] は、完全な音声文字起こしがあり、それに対する実行結果だけを生成したい場合に使用します。話者が話し終えたタイミングを検出する必要がない場合に便利です。たとえば、事前録音された音声がある場合や、ユーザーが話し終えたことが明確なプッシュ・トゥ・トークアプリなどです。
|
||||
2. [`StreamedAudioInput`][agents.voice.input.StreamedAudioInput] は、ユーザーが話し終えたタイミングを検出する必要がある可能性がある場合に使用します。検出された音声チャンクを送信でき、音声パイプラインは「アクティビティ検出」と呼ばれるプロセスを通じて、適切なタイミングでエージェントワークフローを自動的に実行します。
|
||||
|
||||
## 結果
|
||||
## 実行結果
|
||||
|
||||
VoicePipeline 実行の結果は [`StreamedAudioResult`][agents.voice.result.StreamedAudioResult] です。これは、イベントが発生するたびにストリーミングで受け取れるオブジェクトです。いくつかの種類の [`VoiceStreamEvent`][agents.voice.events.VoiceStreamEvent] があります。
|
||||
音声パイプライン実行の結果は [`StreamedAudioResult`][agents.voice.result.StreamedAudioResult] です。これは、発生したイベントをストリーミングできるオブジェクトです。[`VoiceStreamEvent`][agents.voice.events.VoiceStreamEvent] には、次のようないくつかの種類があります。
|
||||
|
||||
1. [`VoiceStreamEventAudio`][agents.voice.events.VoiceStreamEventAudio]:音声チャンクを含みます。
|
||||
2. [`VoiceStreamEventLifecycle`][agents.voice.events.VoiceStreamEventLifecycle]:ターンの開始や終了など、ライフサイクルイベントを通知します。
|
||||
3. [`VoiceStreamEventError`][agents.voice.events.VoiceStreamEventError]:エラーイベントです。
|
||||
1. [`VoiceStreamEventAudio`][agents.voice.events.VoiceStreamEventAudio]。音声チャンクを含みます。
|
||||
2. [`VoiceStreamEventLifecycle`][agents.voice.events.VoiceStreamEventLifecycle]。ターンの開始や終了などのライフサイクルイベントを通知します。
|
||||
3. [`VoiceStreamEventError`][agents.voice.events.VoiceStreamEventError]。エラーイベントです。
|
||||
|
||||
```python
|
||||
|
||||
@@ -63,7 +67,7 @@ async for event in result.stream():
|
||||
# play audio
|
||||
elif event.type == "voice_stream_event_lifecycle":
|
||||
# lifecycle
|
||||
elif event.type == "voice_stream_event_error"
|
||||
elif event.type == "voice_stream_event_error":
|
||||
# error
|
||||
...
|
||||
```
|
||||
@@ -72,4 +76,4 @@ async for event in result.stream():
|
||||
|
||||
### 割り込み
|
||||
|
||||
Agents SDK は現在、[`StreamedAudioInput`][agents.voice.input.StreamedAudioInput] に対する組み込みの割り込みサポートを提供していません。そのため、検出された各ターンごとにワークフローの個別実行がトリガーされます。アプリケーション内で割り込みを処理したい場合は、[`VoiceStreamEventLifecycle`][agents.voice.events.VoiceStreamEventLifecycle] イベントをリッスンできます。`turn_started` は新しいターンが文字起こしされ、処理が開始されたことを示します。`turn_ended` は該当ターンのすべての音声が送信された後にトリガーされます。これらのイベントを利用して、モデルがターンを開始した際に話者のマイクをミュートし、ターンに関連するすべての音声を送信し終えた後にアンミュートする、といった制御が可能です。
|
||||
Agents SDK は現在、[`StreamedAudioInput`][agents.voice.input.StreamedAudioInput] に対する組み込みの割り込み処理を提供していません。代わりに、検出された各ターンごとに、ワークフローの個別の実行がトリガーされます。アプリケーション内で割り込みを処理したい場合は、[`VoiceStreamEventLifecycle`][agents.voice.events.VoiceStreamEventLifecycle] イベントをリッスンできます。`turn_started` は、新しいターンが文字起こしされ、処理が開始されたことを示します。`turn_ended` は、該当するターンのすべての音声がディスパッチされた後にトリガーされます。これらのイベントを使用して、モデルがターンを開始したときに話者のマイクをミュートし、そのターンに関連するすべての音声をフラッシュした後にミュートを解除できます。
|
||||
+18
-14
@@ -1,20 +1,24 @@
|
||||
---
|
||||
search:
|
||||
exclude: true
|
||||
---
|
||||
# クイックスタート
|
||||
|
||||
## 前提条件
|
||||
|
||||
Agents SDK の[クイックスタート手順](../quickstart.md)に従い、仮想環境をセットアップしてください。その後、SDK からオプションの音声依存関係をインストールします。
|
||||
Agents SDK の基本の [クイックスタート手順](../quickstart.md) に従い、仮想環境をセットアップしていることを確認してください。次に、SDK から任意の音声依存関係をインストールします。
|
||||
|
||||
```bash
|
||||
pip install 'openai-agents[voice]'
|
||||
```
|
||||
|
||||
## コンセプト
|
||||
## 概念
|
||||
|
||||
主なコンセプトは [`VoicePipeline`][agents.voice.pipeline.VoicePipeline] です。これは 3 ステップのプロセスです:
|
||||
知っておくべき主な概念は [`VoicePipeline`][agents.voice.pipeline.VoicePipeline] です。これは 3 ステップのプロセスです。
|
||||
|
||||
1. 音声をテキストに変換する音声認識モデル(speech-to-text)を実行します。
|
||||
2. 通常はエージェント的なワークフローであるあなたのコードを実行し、結果を生成します。
|
||||
3. 結果のテキストを音声に戻す音声合成モデル(text-to-speech)を実行します。
|
||||
1. 音声認識モデルを実行して、音声をテキストに変換します。
|
||||
2. 通常はエージェント的なワークフローであるコードを実行して、結果を生成します。
|
||||
3. テキスト読み上げモデルを実行して、結果のテキストを音声に戻します。
|
||||
|
||||
```mermaid
|
||||
graph LR
|
||||
@@ -44,7 +48,7 @@ graph LR
|
||||
|
||||
## エージェント
|
||||
|
||||
まず、いくつかのエージェントをセットアップしましょう。この SDK でエージェントを作成したことがあれば、馴染みがあるはずです。ここでは、複数のエージェント、ハンドオフ、ツールを用意します。
|
||||
まず、いくつかのエージェントをセットアップしましょう。この SDK でエージェントを構築したことがあれば、なじみのある内容です。ここでは、複数のエージェント、ハンドオフ、ツールを用意します。
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
@@ -72,7 +76,7 @@ spanish_agent = Agent(
|
||||
instructions=prompt_with_handoff_instructions(
|
||||
"You're speaking to a human, so be polite and concise. Speak in Spanish.",
|
||||
),
|
||||
model="gpt-4o-mini",
|
||||
model="gpt-5.5",
|
||||
)
|
||||
|
||||
agent = Agent(
|
||||
@@ -80,7 +84,7 @@ agent = Agent(
|
||||
instructions=prompt_with_handoff_instructions(
|
||||
"You're speaking to a human, so be polite and concise. If the user speaks in Spanish, handoff to the spanish agent.",
|
||||
),
|
||||
model="gpt-4o-mini",
|
||||
model="gpt-5.5",
|
||||
handoffs=[spanish_agent],
|
||||
tools=[get_weather],
|
||||
)
|
||||
@@ -88,7 +92,7 @@ agent = Agent(
|
||||
|
||||
## 音声パイプライン
|
||||
|
||||
[`SingleAgentVoiceWorkflow`][agents.voice.workflow.SingleAgentVoiceWorkflow] をワークフローとして使い、シンプルな音声パイプラインをセットアップします。
|
||||
ワークフローとして [`SingleAgentVoiceWorkflow`][agents.voice.workflow.SingleAgentVoiceWorkflow] を使用し、シンプルな音声パイプラインをセットアップします。
|
||||
|
||||
```python
|
||||
from agents.voice import SingleAgentVoiceWorkflow, VoicePipeline
|
||||
@@ -120,7 +124,7 @@ async for event in result.stream():
|
||||
|
||||
```
|
||||
|
||||
## すべてをまとめる
|
||||
## 全体の統合
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
@@ -156,7 +160,7 @@ spanish_agent = Agent(
|
||||
instructions=prompt_with_handoff_instructions(
|
||||
"You're speaking to a human, so be polite and concise. Speak in Spanish.",
|
||||
),
|
||||
model="gpt-4o-mini",
|
||||
model="gpt-5.5",
|
||||
)
|
||||
|
||||
agent = Agent(
|
||||
@@ -164,7 +168,7 @@ agent = Agent(
|
||||
instructions=prompt_with_handoff_instructions(
|
||||
"You're speaking to a human, so be polite and concise. If the user speaks in Spanish, handoff to the spanish agent.",
|
||||
),
|
||||
model="gpt-4o-mini",
|
||||
model="gpt-5.5",
|
||||
handoffs=[spanish_agent],
|
||||
tools=[get_weather],
|
||||
)
|
||||
@@ -191,4 +195,4 @@ if __name__ == "__main__":
|
||||
asyncio.run(main())
|
||||
```
|
||||
|
||||
この例を実行すると、エージェントがあなたに話しかけます のコード例をチェックして、実際にエージェントと会話できるデモをご覧ください。
|
||||
この例を実行すると、エージェントがあなたに話しかけます!エージェントに自分で話しかけられるデモについては、[examples/voice/static](https://github.com/openai/openai-agents-python/tree/main/examples/voice/static) の例を確認してください。
|
||||
@@ -1,14 +1,18 @@
|
||||
---
|
||||
search:
|
||||
exclude: true
|
||||
---
|
||||
# トレーシング
|
||||
|
||||
[エージェントのトレーシング](../tracing.md)と同様に、音声パイプラインも自動的にトレーシングされます。
|
||||
|
||||
基本的なトレーシング情報については上記のトレーシングドキュメントをご参照いただけますが、さらに [`VoicePipelineConfig`][agents.voice.pipeline_config.VoicePipelineConfig] を使ってパイプラインのトレーシングを設定することも可能です。
|
||||
基本的なトレーシング情報については上記のトレーシングドキュメントを参照できますが、さらに [`VoicePipelineConfig`][agents.voice.pipeline_config.VoicePipelineConfig] を通じてパイプラインのトレーシングを設定できます。
|
||||
|
||||
主なトレーシング関連フィールドは以下の通りです:
|
||||
トレーシングに関連する主なフィールドは次のとおりです。
|
||||
|
||||
- [`tracing_disabled`][agents.voice.pipeline_config.VoicePipelineConfig.tracing_disabled]:トレーシングを無効にするかどうかを制御します。デフォルトではトレーシングは有効です。
|
||||
- [`trace_include_sensitive_data`][agents.voice.pipeline_config.VoicePipelineConfig.trace_include_sensitive_data]:トレースに音声の書き起こしなど、機密性の高いデータを含めるかどうかを制御します。これは特に音声パイプライン用であり、Workflow 内部で発生する内容には適用されません。
|
||||
- [`trace_include_sensitive_audio_data`][agents.voice.pipeline_config.VoicePipelineConfig.trace_include_sensitive_audio_data]:トレースに音声データを含めるかどうかを制御します。
|
||||
- [`workflow_name`][agents.voice.pipeline_config.VoicePipelineConfig.workflow_name]:トレースワークフローの名前です。
|
||||
- [`group_id`][agents.voice.pipeline_config.VoicePipelineConfig.group_id]:トレースの `group_id` で、複数のトレースをリンクすることができます。
|
||||
- [`trace_metadata`][agents.voice.pipeline_config.VoicePipelineConfig.tracing_disabled]:トレースに追加するメタデータです。
|
||||
- [`tracing_disabled`][agents.voice.pipeline_config.VoicePipelineConfig.tracing_disabled]: トレーシングを無効にするかどうかを制御します。デフォルトでは、トレーシングは有効です。
|
||||
- [`trace_include_sensitive_data`][agents.voice.pipeline_config.VoicePipelineConfig.trace_include_sensitive_data]: トレースに音声文字起こしなど、潜在的に機微なデータを含めるかどうかを制御します。これは音声パイプライン専用であり、Workflow 内で行われる処理には適用されません。
|
||||
- [`trace_include_sensitive_audio_data`][agents.voice.pipeline_config.VoicePipelineConfig.trace_include_sensitive_audio_data]: トレースに音声データを含めるかどうかを制御します。
|
||||
- [`workflow_name`][agents.voice.pipeline_config.VoicePipelineConfig.workflow_name]: トレースワークフローの名前です。
|
||||
- [`group_id`][agents.voice.pipeline_config.VoicePipelineConfig.group_id]: トレースの `group_id` で、複数のトレースを関連付けることができます。
|
||||
- [`trace_metadata`][agents.voice.pipeline_config.VoicePipelineConfig.trace_metadata]: トレースに含める追加のメタデータです。
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user